LLM Systems & Agentic Engineering Wiki

A deep-dive technical reference on tokenizer internals, decoder transformers, autoregressive pre-training, and autonomous agent loops.

1. Tokenization & Byte-Pair Encoding (BPE)

Tokenization is the discrete translation layer between raw UTF-8 byte sequences and the integer embeddings fed into transformer models. Modern LLMs operate on Byte-Level Byte-Pair Encoding (BPE), popularized in systems like GPT-2, GPT-4 (tiktoken / cl100k_base), and Andrej Karpathy's minbpe.

UTF-8 Byte Fallback

Treats raw text as a sequence of bytes (0–255). This eliminates out-of-vocabulary (OOV) tokens because every character can be decomposed into UTF-8 byte streams.

Iterative Merge Tables

Starts with a base 256-byte vocabulary and repeatedly counts adjacent pair frequencies, merging the most frequent pairs into new composite tokens until reaching target vocabulary size $V$.

Iterative BPE Merge Algorithm

The core training loop iteratively constructs a merge dictionary mapping (token_a, token_b) -> new_token_id:

def get_stats(ids):
    counts = {}
    for pair in zip(ids, ids[1:]):
        counts[pair] = counts.get(pair, 0) + 1
    return counts

def merge(ids, pair, idx):
    newids = []
    i = 0
    while i < len(ids):
        if i < len(ids) - 1 and ids[i] == pair[0] and ids[i+1] == pair[1]:
            newids.append(idx)
            i += 2
        else:
            newids.append(ids[i])
            i += 1
    return newids

Regex Pre-Splitting Patterns

Direct BPE merges across punctuation or spaces can cause semantic fragmentation (e.g., grouping punctuation with words like "world!" as a distinct entity from "world"). Modern tokenizers enforce regex boundaries to prevent cross-category merges.

Tokenizer Standard Regex Splitting Pattern Key Characteristic
GPT-2 's|'t|'re|'ve|'m|'ll|'d| ?\p{L}+| ?\p{N}+| ?[^\s\p{L}\p{N}]+|\s+ Splits contractions, letters, numbers, and punctuation into distinct clusters before BPE.
GPT-4 (cl100k) '(?i:[sdmt]|ll|ve|re)|[^\r\n\p{L}\p{N}]?+\p{L}+|\p{N}{1,3}| ?[^\s\p{L}\p{N}]++[\r\n]*|\s*[\r\n]|\s++$|\s++ Limits digit chunking to 3 digits max (optimizes math representations) and isolates whitespace.

Special Tokens & Vocabulary Construction

Special delimiter tokens provide out-of-band structural signaling to the model without conflicting with raw text:

2. Decoder-Only Transformer Architecture

Following the modern standard (Llama, Mistral, nanoGPT, GPT-4), modern autoregressive models use a decoder-only layout with pre-normalization and Rotary Position Embeddings (RoPE).

RMSNorm Pre-Normalization

Replaces traditional LayerNorm by normalizing inputs via Root Mean Square without mean-centering, reducing computational latency across layer transitions.

SwiGLU Activation

Replaces ReLU/GELU with a Swish-Gated Linear Unit in the MLP block, yielding higher empirical validation efficiency per parameter.

Decoder Forward Pass Overview

x = x + attention(rmsnorm(x))  # Self-Attention with causal lower-triangular mask
x = x + mlp(rmsnorm(x))        # Gated feed-forward network

Attention Optimizations: RoPE & KV-Cache

In autoregressive generation, past token representations are cached to prevent $O(N^2)$ recalculations during sequential token emission:

Pre-training Dynamics & Optimization

Autoregressive training minimizes Cross-Entropy loss over token sequence predictions:

$$\mathcal{L} = -\sum_{i=1}^{T} \log P(x_i \mid x_1, x_2, \dots, x_{i-1})$$

3. Agentic Systems & Reasoning Loops

An AI Agent extends an autoregressive language model into a stateful, interactive environment through continuous perception-action-observation cycles.

ReAct Framework

Interleaves Thought (chain-of-thought reasoning), Action (tool invocation), and Observation (environment feedback) before outputting final answers.

Deterministic Control Loops

Wraps the LLM in an execution loop that enforces schema validation, handles runtime tool errors, and maintains memory buffers across iterative steps.

Canonical ReAct Execution Pattern

Thought: I need to check the server firewall status before recommending changes.
Action: execute_bash({"command": "sudo ufw status numbered"})
Observation: Status: active \n [1] 22/tcp ALLOW IN Anywhere
Thought: Port 22 is open, but Port 443 is missing. I must open 443.
Action: execute_bash({"command": "sudo ufw allow 443/tcp"})
Observation: Rule added
Thought: The port has been successfully authorized. I will now inform the user.
Final Answer: Port 443/tcp has been successfully opened in UFW.

Tool Invocation & Structured Output Contracts

Modern models are instruction-tuned to emit constrained JSON payloads targeting predefined function signatures:

Post-Training: SFT, DPO & Reinforcement

Transforming a base pre-trained completion model into a capable instruction-following agent:

Training Stage Objective Primary Methodologies
Supervised Fine-Tuning (SFT) Teaches conversational turn format and deterministic tool syntax formatting. Multi-turn ChatML prompt masking, loss calculated exclusively on assistant tokens.
Direct Preference Optimization (DPO) Aligns outputs to human/evaluator preferences without separate reward models. Implicit reward optimization using paired preference datasets (chosen vs. rejected responses).
RL with Verifiable Rewards (RLVR / GRPO) Incentivizes multi-step reasoning, coding, and mathematical verification. Group Relative Policy Optimization (GRPO) evaluating rule-based test pass rates and compiler checks.

4. Curated Sources & Community Research