Case study · B2B outbound
Limelight
Cold email that books qualified meetings for SaaS founders.
Limelight runs outbound end to end for B2B SaaS founders. They build the prospect list, write the copy, and run the campaigns that put qualified sales conversations on the calendar.
Why enrichment is the constraint
List quality is the product. Every prospect has to be classified and qualified before a single email goes out, so the cost of enrichment caps how large a list they can afford to build.
Six weeks on ZeroGPU
Jul 29 – Sep 6, 2026
10.97B in · 667M out
Sat Aug 29
Zero failed requests
The problem
Before ZeroGPU: a cost wall at millions of rows
"Flow before was using Claude Haiku, then I switched to Gemini Flash-Lite. But it was still quite expensive because some enrichments can be millions of rows."
The old pipeline
Scrape prospect website
raw homepage text · ~3,800 tokens/row
Frontier LLM classifies
Claude Haiku → Gemini 3.5 Flash-Lite
Extract client-specific signals
same model, second pass
Qualified list
gated by budget, not by ambition
Cost for the same 11.63B tokens
Input-heavy work: ~3,813 tokens in per row means cost scales with list size.
Published list prices · September 2026
The result
After ZeroGPU: the same job, 9.2× cheaper
Same 10.97B input / 667M output tokens
vs Gemini 3.5 Flash-Lite: 3.2× cheaper, $3,396 saved.
What they run now
gpt-oss-120b reads scraped site text and, in one pass:
- →Classifies the company into a specific type
- →Extracts client-specific signals off the homepage
Validated on a 100–150 row labelled set at 90% precision and recall.
Source: ZeroGPU Analytics Engine · org dab6d0f7 · Jul 29 – Sep 6, 2026
What is ZeroGPU
Not every task needs a frontier model
ZeroGPU runs right-sized open and purpose-built models behind one OpenAI-compatible API, so high-volume work stops being priced like frontier reasoning.
Right-sized models
Pick the smallest model that clears your accuracy bar, from 86M-parameter classifiers to 120B reasoning models.
One API, no GPU ops
OpenAI-compatible endpoints. No clusters to provision, no autoscaling to babysit, no idle GPUs to pay for.
Pay per token
Purpose-built models from $0.02 / M input tokens. Costs track usage, not reserved capacity.
What that looked like for Limelight
Self-serve, no sales call
81,482 requests in one hour
Across 2.88M calls
Down from $4.97
zerogpu.ai
For outreach agencies
The same stack, for every step of your pipeline
An example enrichment flow
Domain in
raw list
Classify
company type, ICP fit
Extract
signals to typed JSON
Scrub PII
compliance before send
Personalise
reasoning model
Models built for this work
| Model | What it does for outbound |
|---|---|
| gliner2-base-v1 | 205MPull typed JSON fields straight out of messy page text |
| deberta-v3-small | 142MScore ICP fit and intent against your own labels |
| gliner-multi-pii-v1 | 300MDetect and redact PII before a list leaves your systems |
| llama-3.1-8b-instruct-fast | 8BSummarise long pages and threads at 128K context |
| qwen3-30b-a3b-fp8 | 30.5BDraft personalised opening lines across an entire list |
| deepseek-v4-flash | 284BRead a company's whole site in one call before you write |
| gpt-oss-120b | 120BClassify, extract and personalise, Limelight's workhorse |
hello@zerogpu.ai