Route 85% of traffic to a small model you fine-tuned.

A Llama 3.2 3B specialist, trained on your own documents with Unsloth QLoRA and served by vLLM, handles routine extraction on a single GPU. A LiteLLM gateway escalates only the hard cases to the frontier model.

Local
85.0%
Escalated
15.0%
Routed
1,000
Live routing sketch. The 85 / 15 split is the design target, not a measured production figure.

Routine extraction does not need a frontier model.

Relying exclusively on proprietary frontier models such as GPT-4o for routine data extraction creates massive API bills and data privacy issues. Every invoice, contract and form leaves your boundary, and you pay frontier prices to read it.

  • Unsloth QLoRA
  • Llama 3.2 3B
  • vLLM
  • LiteLLM Proxy
  • FastAPI
  • NVIDIA L4 GPU

Move the sliders. Watch the invoice change.

A blueprint simulator. Every figure here is modelled, not measured from a live deployment.

1,000,000
85%

Model: direct = volume x $0.015. Hybrid = $260 L4 GPU + (1 - rate) x volume x $0.015 + serving overhead. The overhead is set to $0 because the $260 GPU line already covers serving, which makes the default (1M requests, 85%) land on exactly $2,510 per month.

Direct, 100% GPT-4o$15,000
Hybrid, SLM + escalations$2,510
L4 GPU $260Escalations $2,250
Net saving per month $12,490
83.3%lower
p95 latency, benchmark claim 1,400 ms ~320 ms

From raw documents to a routed request.

  1. 1

    Prepare the domain dataset

    Raw documents become instruction pairs, then split for instruction tuning so the model is graded on examples it never trained on.

  2. 2

    Fine-tune with 4-bit QLoRA

    Unsloth quantises Llama 3.2 3B to 4-bit and trains rank-16 adapters, small enough to iterate on a single NVIDIA L4.

  3. 3

    Merge adapters and serve

    LoRA adapters are merged and exported to a vLLM container, with PagedAttention keeping GPU memory efficient under concurrent load.

  4. 4

    Evaluate complexity and route

    The LiteLLM gateway scores each request for complexity and routes it semantically, with confidence scores deciding when to fall back.

  5. 5

    Run local, escalate the rest

    85% of requests execute on the local SLM. The remaining 15% escalate automatically to the frontier model.

Frontier models are brilliant, and wildly overpriced for reading the same kinds of document all day. A small model trained on your own data does the repetitive work locally, on one GPU, behind your firewall. Only the genuinely hard cases ever leave the building.

Four things you can run on day one.

Hover a card. Code shown is illustrative.

Unsloth QLoRA training scripts

Reproducible fine-tuning with hyperparameter configs checked into the repo.

model = FastLanguageModel.from_pretrained(
  "unsloth/Llama-3.2-3B",
  load_in_4bit=True,
)
model = FastLanguageModel.get_peft_model(
  model, r=16,
  target_modules=["q_proj", "v_proj", ...],
)

Data synthesis pipeline

Converts raw text into clean instruction datasets.

vLLM deployment

Containerised, with OpenAI-compatible REST endpoints.

POST /v1/chat/completions

LiteLLM routing proxy

Dynamic routing with confidence-score fallbacks.

Let's cut your inference bill, not your quality.

Fine-tuning, serving and routing, built to run on hardware you control.

Start a conversation