What actually happens when you run AI models locally on your laptop?

What actually happens when you run AI models locally on your laptop?

Executing a query against a cloud service like Claude or GPT-4o feels deceptively simple. You send a payload over TLS, let an enterprise data centre handle the compute across high-bandwidth GPU clusters, and stream back tokens.

Running an open-weight model like DeepSeek on your workstation removes that safety net. The remote infrastructure disappears, and your machine's physical limits — memory bandwidth, unified RAM allocations, and thermal headroom — immediately dictate what is possible.

When you run ollama run deepseek-r1:8b or launch a local llama.cpp server, your machine triggers a specific sequence of system calls, matrix operations, and dynamic memory allocations. Understanding this runtime pipeline reveals why local inference feels instant for brief tasks, stalls when context expands, and remains fundamentally bounded by memory bandwidth rather than GPU core counts.

The reality of local "DeepSeek"

The flagship DeepSeek-R1 model is a 671-billion parameter Mixture-of-Experts (MoE) architecture. At 16-bit floating-point precision (FP16), storing its weights requires over 1.3 terabytes of VRAM — a footprint that demands enterprise multi-GPU nodes.

When running DeepSeek on a local laptop or workstation, you are almost certainly running one of two things:

  • A distilled reasoning model: a dense architecture (such as 8B or 14B parameters, built on Qwen or Llama bases) fine-tuned on reasoning traces from DeepSeek-R1.

  • A compressed MoE export: a heavily quantised GGUF export of the 671B model (such as IQ1_S or Q2_K) mapped across system memory, sacrificing precision to fit consumer RAM.

While parameter counts differ, the underlying local execution path is identical.

Storage to memory: mmap() and zero-copy allocation

Local engines do not perform a standard disk-to-RAM copy of a 5 GB file on startup. Doing so would cause long initialization delays and waste system memory.

Instead, runtimes rely on container formats like GGUF, which align tensor weights into contiguous binary blocks mapped directly to memory boundaries. To load these weights, runtimes call the POSIX system function mmap() (memory mapping).

Rather than reading the file immediately, mmap() maps the file directly into the process's virtual address space. The operating system kernel then loads specific disk pages into physical RAM on demand as the engine accesses individual model layers.

On Apple Silicon architectures, Unified Memory Architecture (UMA) allows the CPU and GPU to read directly from this shared RAM pool without PCIe transfers. On discrete GPU setups, mapped weights must instead travel across the PCIe bus into dedicated VRAM.

Weight compression and dynamic dequantisation

An uncompressed 8-billion parameter model in FP16 precision requires roughly 16 GB of memory (2 bytes per parameter). Because dedicating 16 GB solely to weights is impractical on standard hardware, runtimes rely on quantisation.

In a common Q4_K_M (4-bit medium) scheme:

  • Model weights are divided into discrete blocks (typically 32 or 256 parameters).

  • High-precision floating-point weights are scaled and stored as 4-bit integers.

  • A 32-bit floating-point scale factor is stored alongside each block to reconstruct approximate weight values during execution.

This drops the static footprint of an 8B model to roughly 4.5 GB. The operational trade-off occurs during generation: on every forward pass, the GPU must dynamically dequantise these 4-bit integers back into floating-point numbers in registers before performing matrix operations.

Prefill vs decode: the memory bandwidth bottleneck

Inference runs in two distinct phases, each bounded by entirely different hardware resources.

Prompt processing (prefill)

During prefill, the engine ingests the entire prompt at once. Because all input tokens are known simultaneously, matrix multiplications execute in parallel using optimized BLAS or Metal performance kernels. This phase is compute-bound, saturating GPU floating-point throughput (FLOPS).

Token generation (decode)

Token generation is autoregressive: generating token N requires the output of token N-1. To generate a single token, the engine must stream the entire set of model weights through memory registers once.

Token generation speed is bounded by memory bandwidth, not compute performance.

On a machine with 150 GB/s memory bandwidth running a 4.5 GB model, the theoretical maximum generation speed is bounded by how fast memory can deliver weight data to the processor:

Theoretical Max Throughput = 150 GB/s / 4.5 GB
                            ≈ 33 tokens/second

Even with significantly more GPU compute cores, generation speed cannot exceed the physical rate at which the memory controller streams weight bytes into execution registers.

Dynamic overhead: the KV cache

Beyond static weight storage, running a model requires dynamic working memory. As reasoning models like DeepSeek-R1 generate long chain-of-thought outputs, the Key-Value (KV) cache grows continuously.

To avoid recomputing key and value vectors for past tokens at every step, the engine stores historical attention states in dynamic memory. This memory footprint grows linearly with sequence length:

KV Cache Size = 2 × (Sequence Length) × (Layers)
              × (Attention Heads) × (Head Dimension)
              × (Precision Bytes)

For an 8B model running an 8,192-token context window at 16-bit precision, the KV cache adds 1.0 GB to 1.5 GB of dynamic memory overhead on top of static model weights. If total memory allocation exceeds available physical RAM, the operating system swaps memory pages to disk, causing generation speed to drop dramatically.

Practical deployment considerations

Deploying local models effectively requires setting explicit operational boundaries rather than treating them as drop-in cloud API replacements.

  • Limit context allocations explicitly: configure explicit context caps (e.g. num_ctx 4096) in local runtime configurations to prevent extended reasoning chains from exhausting physical RAM.

  • Align quantisation to task requirements: use 4-bit or 5-bit quantisations ( Q4_K_M, Q5_K_M) for tasks requiring precise logic, syntax, or mathematical reasoning. Reserve aggressive 2-bit or 3-bit quantisations for basic prose generation or text classification.

  • Audit system boundaries: while local runtimes process prompts on-device, surrounding UI or agent frameworks may still route requests to external search APIs, remote embedding services, or remote telemetry endpoints.

Local inference is fundamentally a systems engineering trade-off — balancing physical memory bandwidth, context limits, and precision against local control.

Sources & further reading

More reading