15줄의 JavaScript에서 LLM VRAM 추정: 가중치, KV 캐시 및 헤드룸

작성자

카테고리:

← 피드로
DEV Community · Mineshop · 2026-09-30 개발(SW)

Disclosure: this post is published by Mineshop.eu, an EU hardware shop that sells graphics cards and AI workstations. It was written by an AI agent working for the shop; every number below is reproduced by the tests in the repository linked at the end.

“How much VRAM do I need to run this model locally?” is the first question in almost every local-LLM thread. You don’t need a spreadsheet for a good first answer. Three terms cover most of it: the weights, the KV cache and some headroom for the runtime.

1. Weights: parameters × bits ÷ 8

The raw size of the weights is simply the parameter count times the bits per parameter, divided by 8 to get bytes. We report everything in GiB (2³⁰ bytes), because that is how GPU memory is sized: a “16 GB” card has 16 GiB.

Model 4-bit 16-bit (FP16/BF16) 8B 3.73 GiB 14.90 GiB 14B 6.52 GiB 26.08 GiB 32B 14.90 GiB 59.60 GiB 70B 32.60 GiB 130.39 GiB

Real quantised files are a bit larger than the raw number: formats such as GGUF Q4_K_M store scales and metadata, and some tensors stay at higher precision. We model that as an adjustable overhead (15 % by default). If you already have the model file, its size on disk is a better starting point.

2. KV cache: the part people forget

Every token in the context keeps a key and a value vector per layer and per KV head:

KV bytes = 2 × layers × kv_heads × head_dim × tokens × parallel_sequences × bytes_per_value

Enter fullscreen mode Exit fullscreen mode

For a Llama-3-8B-class model (32 layers, 8 KV heads thanks to grouped-query attention, head dimension 128) with an 8,192-token context and an FP16 cache, that is exactly 2 × 32 × 8 × 128 × 8192 × 2 = 1,073,741,824 bytes, so 1 GiB. It scales linearly: 32K tokens cost 4 GiB, four parallel sequences cost four times as much, and an 8-bit cache halves it (if your engine supports one).

A 70B-class model (80 layers, 8 KV heads, head dimension 128) at 32K tokens needs 10 GiB for the cache alone, on top of the weights.

3. Headroom

The runtime needs memory too: the CUDA context, compute buffers and fragmentation. We use a 2 GiB reserve as a starting assumption, and more if the same GPU also drives your desktop.

The whole thing in JavaScript

function estimate(p) {
  const GiB = 2 ** 30;
  const weights = p.parameters * 1e9 * p.bits / 8 / GiB;
  const quantOverhead = weights * p.overhead / 100;
  const kv = 2 * p.layers * p.kvHeads * p.headDim * p.context * p.parallel * p.kvBytes / GiB;
  return { weights, quantOverhead, kv, reserve: p.reserve,
           total: weights + quantOverhead + kv + p.reserve };
}

const llama8b = { parameters: 8, bits: 4, overhead: 15, layers: 32, kvHeads: 8,
                  headDim: 128, context: 8192, parallel: 1, kvBytes: 2, reserve: 2 };
console.log(estimate(llama8b).total.toFixed(2)); // "7.28"

Enter fullscreen mode Exit fullscreen mode

The tests in the repository pin the behaviour that matters: the KV term is exactly 1 GiB for the example above, doubles with twice the context and halves with an 8-bit cache. The full version also rejects invalid input (NaN, zero sequences) instead of returning a silent number.

const assert = require('node:assert/strict');
assert.equal(estimate(llama8b).kv, 1);
assert.equal(estimate({ ...llama8b, context: 16384 }).kv, 2);
assert.equal(estimate({ ...llama8b, kvBytes: 1 }).kv, 0.5);

Enter fullscreen mode Exit fullscreen mode

What that means for common sizes

4-bit weights, 15 % format overhead, FP16 KV cache, one sequence, 2 GiB reserve:

Model class Context Weights + overhead KV cache Total 8B (32 layers) 8K 4.28 GiB 1.00 GiB 7.28 GiB 8B (32 layers) 32K 4.28 GiB 4.00 GiB 10.28 GiB 14B (40 layers) 8K 7.50 GiB 1.25 GiB 10.75 GiB 32B (64 layers) 8K 17.14 GiB 2.00 GiB 21.14 GiB 70B (80 layers) 8K 37.49 GiB 2.50 GiB 41.99 GiB 70B (80 layers) 32K 37.49 GiB 10.00 GiB 49.49 GiB

So in practice:

  • 16 GB cards (RTX 5060 Ti 16 GB, RTX 5070 Ti, RTX 5080) run 8–14B models at 4-bit with room for a normal context. A 32B model does not fit; part of it would have to be offloaded to system RAM, which is much slower.
  • 24–32 GB is where 32B models become comfortable.
  • 70B needs about 42 GiB at 8K context, which means two cards or a 48 GB+ professional GPU; with long context, 96 GB cards such as the RTX PRO 6000 Blackwell leave real headroom.

Where the estimate breaks

This is a planning number for dense transformers with a full attention cache, not a benchmark. Mixture-of-experts models need memory for all experts, not just the active ones. Sliding-window attention, MLA (DeepSeek-style) and hybrid architectures change the KV term a lot. Vision encoders and engine-specific buffers come on top. And two 24 GB cards are not one 48 GB pool: the engine, layer split and per-GPU buffers decide what actually fits.

Try it

We put the same formula into a small browser tool with presets, a custom-architecture panel and CSV export. It runs offline, with no cookies or analytics: LLM Memory Planner (source and tests on GitHub). There are also German, French and Latvian editions.

If you’re pricing hardware after running the numbers: graphics cards and AI workstations on Mineshop.eu (see the disclosure above). Corrections to the formulas are very welcome in the comments.

원문에서 계속 ↗