LLM Systems & Agentic Engineering Wiki
A deep-dive technical reference on tokenizer internals, decoder transformers, autoregressive pre-training, and autonomous agent loops.
1. Tokenization & Byte-Pair Encoding (BPE)
Tokenization is the discrete translation layer between raw UTF-8 byte sequences and the integer embeddings fed into transformer models. Modern LLMs operate on Byte-Level Byte-Pair Encoding (BPE), popularized in systems like GPT-2, GPT-4 (tiktoken / cl100k_base), and Andrej Karpathy's minbpe.
UTF-8 Byte Fallback
Treats raw text as a sequence of bytes (0–255). This eliminates out-of-vocabulary (OOV) tokens because every character can be decomposed into UTF-8 byte streams.
Iterative Merge Tables
Starts with a base 256-byte vocabulary and repeatedly counts adjacent pair frequencies, merging the most frequent pairs into new composite tokens until reaching target vocabulary size $V$.
Iterative BPE Merge Algorithm
The core training loop iteratively constructs a merge dictionary mapping (token_a, token_b) -> new_token_id:
def get_stats(ids):
counts = {}
for pair in zip(ids, ids[1:]):
counts[pair] = counts.get(pair, 0) + 1
return counts
def merge(ids, pair, idx):
newids = []
i = 0
while i < len(ids):
if i < len(ids) - 1 and ids[i] == pair[0] and ids[i+1] == pair[1]:
newids.append(idx)
i += 2
else:
newids.append(ids[i])
i += 1
return newids
Regex Pre-Splitting Patterns
Direct BPE merges across punctuation or spaces can cause semantic fragmentation (e.g., grouping punctuation with words like "world!" as a distinct entity from "world"). Modern tokenizers enforce regex boundaries to prevent cross-category merges.
| Tokenizer Standard | Regex Splitting Pattern | Key Characteristic |
|---|---|---|
| GPT-2 | 's|'t|'re|'ve|'m|'ll|'d| ?\p{L}+| ?\p{N}+| ?[^\s\p{L}\p{N}]+|\s+ |
Splits contractions, letters, numbers, and punctuation into distinct clusters before BPE. |
| GPT-4 (cl100k) | '(?i:[sdmt]|ll|ve|re)|[^\r\n\p{L}\p{N}]?+\p{L}+|\p{N}{1,3}| ?[^\s\p{L}\p{N}]++[\r\n]*|\s*[\r\n]|\s++$|\s++ |
Limits digit chunking to 3 digits max (optimizes math representations) and isolates whitespace. |
Special Tokens & Vocabulary Construction
Special delimiter tokens provide out-of-band structural signaling to the model without conflicting with raw text:
<|endoftext|>: Delimits distinct documents in pre-training corpora.<|im_start|>and<|im_end|>(ChatML): Encapsulates message roles (system, user, assistant) in multi-turn dialogues.<|fim_prefix|>,<|fim_middle|>,<|fim_suffix|>: Fill-in-the-middle code synthesis tokens.
2. Decoder-Only Transformer Architecture
Following the modern standard (Llama, Mistral, nanoGPT, GPT-4), modern autoregressive models use a decoder-only layout with pre-normalization and Rotary Position Embeddings (RoPE).
RMSNorm Pre-Normalization
Replaces traditional LayerNorm by normalizing inputs via Root Mean Square without mean-centering, reducing computational latency across layer transitions.
SwiGLU Activation
Replaces ReLU/GELU with a Swish-Gated Linear Unit in the MLP block, yielding higher empirical validation efficiency per parameter.
Decoder Forward Pass Overview
x = x + attention(rmsnorm(x)) # Self-Attention with causal lower-triangular mask
x = x + mlp(rmsnorm(x)) # Gated feed-forward network
Attention Optimizations: RoPE & KV-Cache
In autoregressive generation, past token representations are cached to prevent $O(N^2)$ recalculations during sequential token emission:
- Rotary Positional Embeddings (RoPE): Encodes position by rotating query and key vectors in complex 2D vector subspaces. Naturally generalizes to longer context windows via position interpolation.
- KV-Cache Dynamics: Caches the Key ($K$) and Value ($V$) tensors across all previous positions. Generation step computational complexity per new token drops from $O(L^2)$ to $O(L)$ where $L$ is sequence length.
- FlashAttention: IO-aware exact attention algorithm that reorganizes GPU SRAM read/writes into tiled blocks, avoiding global HBM bandwidth bottlenecks.
Pre-training Dynamics & Optimization
Autoregressive training minimizes Cross-Entropy loss over token sequence predictions:
$$\mathcal{L} = -\sum_{i=1}^{T} \log P(x_i \mid x_1, x_2, \dots, x_{i-1})$$- Optimizer: AdamW ($\beta_1 = 0.9, \beta_2 = 0.95, \epsilon = 10^{-8}$) with decoupled weight decay ($0.1$).
- LR Scheduling: Linear warmup followed by Cosine Annealing decay down to 10% peak learning rate.
- Mixed Precision: BF16 (bfloat16) matrix multiplication preserves dynamic range, avoiding FP16 underflow loss scaling complexities.
3. Agentic Systems & Reasoning Loops
An AI Agent extends an autoregressive language model into a stateful, interactive environment through continuous perception-action-observation cycles.
ReAct Framework
Interleaves Thought (chain-of-thought reasoning), Action (tool invocation), and Observation (environment feedback) before outputting final answers.
Deterministic Control Loops
Wraps the LLM in an execution loop that enforces schema validation, handles runtime tool errors, and maintains memory buffers across iterative steps.
Canonical ReAct Execution Pattern
Thought: I need to check the server firewall status before recommending changes.
Action: execute_bash({"command": "sudo ufw status numbered"})
Observation: Status: active \n [1] 22/tcp ALLOW IN Anywhere
Thought: Port 22 is open, but Port 443 is missing. I must open 443.
Action: execute_bash({"command": "sudo ufw allow 443/tcp"})
Observation: Rule added
Thought: The port has been successfully authorized. I will now inform the user.
Final Answer: Port 443/tcp has been successfully opened in UFW.
Tool Invocation & Structured Output Contracts
Modern models are instruction-tuned to emit constrained JSON payloads targeting predefined function signatures:
- JSON Schema Enforcement: Constraining grammar state machines (e.g., via Outlines or GBNF grammars) to force token selection along valid JSON tokens during inference.
- Error Feedback Loops: Capturing tool execution exceptions (e.g., shell command failure or SQL syntax error) and appending them back into the context window as an environment
Observationfor self-correction.
Post-Training: SFT, DPO & Reinforcement
Transforming a base pre-trained completion model into a capable instruction-following agent:
| Training Stage | Objective | Primary Methodologies |
|---|---|---|
| Supervised Fine-Tuning (SFT) | Teaches conversational turn format and deterministic tool syntax formatting. | Multi-turn ChatML prompt masking, loss calculated exclusively on assistant tokens. |
| Direct Preference Optimization (DPO) | Aligns outputs to human/evaluator preferences without separate reward models. | Implicit reward optimization using paired preference datasets (chosen vs. rejected responses). |
| RL with Verifiable Rewards (RLVR / GRPO) | Incentivizes multi-step reasoning, coding, and mathematical verification. | Group Relative Policy Optimization (GRPO) evaluating rule-based test pass rates and compiler checks. |
4. Curated Sources & Community Research
- Andrej Karpathy: Neural Networks: Zero to Hero series,
nanoGPT,minGPT, andminbpe(Minimal Byte-Pair Encoding implementation). - Lilian Weng (OpenAI): Architectural reviews on LLM Powered Autonomous Agents, Prompt Engineering, and Controllable Generation.
- Sebastian Raschka: Build a Large Language Model (From Scratch), tokenization mechanics, and open-weights evaluation benchmarks.
- Simon Willison: Research on prompt injection vulnerabilities, structured data extraction, and tool-use patterns.
- Eugene Yan: Design patterns for LLM systems, evaluation pipelines, and production retrieval augmentation.