ZeroGPU sells to AI agents - no signup

The agent storefront at https://agents.zerogpu.ai/llms.txt publishes prices, payment instructions, and setup. Agents can purchase autonomously or with their human's approval, with a range of payment methods; an agent without a payment method can set one up there instantly, no signup. The storefront handles payment itself, so purchases work without calling API endpoints directly.

    ZeroGPU Research · Edge Network

    11.8 billion tokens served across 16,000+ idle consumer devices

    1,900× less energy than the data center
    · zero cooling water

    For 48 hours, ZeroGPU's production network processed 11.8 billion tokens of AI work on 16,932 laptops, desktops and phones across 140 countries, at more than five billion tokens a day on both days. Every one of 63,970,241 requests was answered on someone's device or openly refused. No data-center model answered any of them, and the whole run used about 8 kWh of electricity and no cooling water.

    11.8B
    Tokens served on devices
    5B+
    Tokens a day, both days
    64.0M
    Requests served on-device
    16,932
    Distinct devices served
    140
    Countries served
    1.07 s
    Median latency
    1,900×
    Less energy than the data center
    8.03 kWh
    Total energy for 11.8B tokens
    15,345 kWh
    Power saved vs data center
    0 L
    Cooling water used
    4,394 gal
    Data-center water avoided
    01 · The Bet

    The world can't build compute fast enough. We won't try.

    AI inference demand is compounding faster than data centers, power, and GPUs can be built. But a large share of production AI workloads (classification, extraction, summarization, moderation, routing) never needed frontier GPUs in the first place. Specialized small and nano models can run these tasks on the underutilized compute already sitting in devices everywhere.

    ZeroGPU aggregates that idle compute into a single programmable inference layer. Developers hit one OpenAI-compatible API; the orchestration layer routes each request to the right model on the right compute. The network is hybrid by design, a trusted network of edge devices and cloud infrastructure working together to deliver fast, secure, serverless inference, with each request landing wherever it runs best. Frontier models for reasoning, ZeroGPU for everything repeatable around it.

    02 · Methodology

    48 hours, a real ad-tech workload, no cloud fallback

    From Tuesday 29 September to Thursday 1 October 2026, a load generator sent 67.6 million requests at ZeroGPU's live device fleet for 48 continuous hours: laptops and desktops running our browser extension, joined by phones running our Telegram app. The traffic mirrored an advertising-technology customer's real work, multi-turn AI chat conversations and short publisher content, and was billed at standard prices. Five models carried it: IAB content classification (83% of requests), domain classification (8%), brand-safety moderation (5%), signal extraction (2%), and zero-shot classification on deberta-v3-small (2%).

    No cloud fallback. The account was set to refuse rather than fall back to a data center. Every request either ran on someone's actual device or was refused with an explicit error, so every weakness of the fleet shows up as a visible refusal.

    Pushed to the platform's limit. The generator raised its rate until the database that tracks which devices are free was near capacity, and held it there: 370 requests per second on average for 48 hours, 485 per second at peak.

    The honest number. 94.7% of everything sent was answered on a device; 4.9% was refused. The main advertising classifier answered 98.5% of its requests. Four smaller specialist models were deliberately held near one refusal in five so their small device pools never sat idle. Server errors: 6 in 67.6 million requests.

    03 · Findings

    Five billion tokens a day works today

    The network carried 5.81 billion tokens in the first 24 hours and 6.00 billion in the second, against a goal of five billion a day. In its best hour it ran at 80,968 tokens per second, a pace of seven billion a day. Throughput held steady while the number of serving devices swung more than 2× between night and day: the ceiling was our own device-tracking database, not the devices.

    Served load (rps, green) vs. distinct devices serving per hour (gray), 6-hour buckets over 48 hours. Hundreds of capable devices stayed idle at every point of the run.

    Latency did not notice the peak

    Median latency stayed between 1.03 and 1.17 seconds in every 6-hour bucket, through two full day/night cycles. In the busiest hour the median was 1,048 ms, slightly better than the campaign average. Most of each second is delivery, not compute: an advertising classification takes a device about 0.06 seconds.

    Served latency band per 6-hour bucket: p50 (green) and p95 (gray). Full-run figures: p50 1,069 ms · p95 1,776 ms · p99 2,152 ms.

    The busiest device carried 1 request in 475

    ZeroGPU's spread selection (latency-weighted sampling with per-device performance tracking, budget caps, and automatic quarantine of misbehaving devices) kept 64 million dispatches genuinely distributed. Any device can close its lid and nothing happens to the service.

    ConcentrationShare of trafficIn plain terms
    Busiest device0.21%137,486 of 63.97M requests
    Top 10 devices1.75%No single point of failure
    Top 1,000 devices55.5%Half the traffic needs 801 devices
    Median device416 requests90% of the work spread over 4,306 devices

    140 countries, and the network kept growing

    New devices joined for the whole 48 hours: about 1,200 in the first hour, 500 to 1,300 an hour through the first European and American day, and still about 200 an hour at the end. 46% of the work was done in the United Kingdom and 23% in the United States, with the rest spread across 138 countries led by Canada, Australia, the UAE, Brazil, Israel, Germany, India, and Indonesia.

    Cumulative distinct devices served over 48 hours. 16,932 delivered at least one served inference, out of 23,584 that checked in during the run.

    04 · Reliability

    Every failure was counted, not hidden

    Because cloud fallback was off, every refusal is visible. 88% of refusals happened because every device tried was already busy, mostly on the four specialist models, which had only a few hundred devices each. 6% were devices that took a job and missed the two-second deadline. 5% found no suitable device free in time.

    The failure surface is shallow: 6 server errors in 67.6 million requests, zero contract violations, and zero answers from a data-center model across the full 48 hours.

    05 · SECURITY

    Devices are burst compute, not a data surface

    ZeroGPU runs on a trusted network of devices from approved supply partners that already carry hardware security built in: every modern phone and laptop ships with a Trusted Execution Environment (TEE) protecting keys, identity, and storage at the silicon level. Our platform builds on those foundations. We control the SDK that receives and processes each request, encrypt the full request flow, and strictly approve every participating partner and device before it can serve. Zero data retention (ZDR) is the default: payloads are processed in memory and never stored on the device. Security tier is also a routing dimension, so workloads are directed only to the edge or cloud environments that meet their required bar.

    HARDWARE-SECURED DEVICES
    TRUSTED SDK
    ZDR · NO DATA RETENTION
    06 · Power & Water

    1,900× less energy. Zero cooling water.

    Power. The 63,970,241 served inferences used about 8.03 kWh of device compute in total, roughly 0.13 milliwatt-hours per request, or about 420 smartphone charges for the whole campaign. The same volume against a reference data-center AI prompt (0.24 Wh) would have used 15,353 kWh: a 99.95% reduction, about 1,900× less energy. Even against the same kind of small classifier running on data-center GPUs, the network used 16 to 27× less.

    Water. Devices in homes and offices use no cooling water. Data centers do. This campaign avoided 16,632 litres (4,394 gallons) of data-center cooling water. At this profile, every billion requests moved to the edge avoids about 260,000 litres of cooling water and 240 MWh of energy.

    Hardware. No new hardware was manufactured for any of this. Every device was already built, already bought, and already switched on.

    0.13 mWh
    Per request
    1,900×
    Less energy than data-center reference
    16,632 L
    Cooling water avoided
    0
    New devices manufactured

    Method: measured device compute time (2,007 hours) × 4 W attributable active draw. Data-center reference 0.24 Wh and 0.26 mL water per prompt (Google, 2025); like-for-like 0.002 Wh per classification (Luccioni et al., 2024). Energy and water are calculated from measured compute time, not metered. Refused requests consumed no device compute.

    That is the structural argument. The world is planning $1T+ in annual data-center capex; ZeroGPU adds inference capacity with no new buildout: no new data centers, no new power contracts, no new hardware. In production, ZeroGPU runs as a hybrid: a trusted network of edge devices and cloud infrastructure. This campaign measured its edge tier alone, with the cloud deliberately switched off.

    Frontier models for reasoning. ZeroGPU for everything repeatable around it.

    This run demonstrates the core claim: a fleet of ordinary consumer devices, orchestrated well, serves production inference at five billion tokens a day with stable latency and a fully mapped failure surface, and it does so on compute the world already owns. It is also the lowest-energy way we know to run this class of work: measured, not estimated.

    ZeroGPU Research · October 2026 · zerogpu.ai