A commercialized secondary market for AI compute has quietly materialized around the arbitrage of promotional credits and venture grants. According to an investigation by Matt Lenhard published on Vectoral on August 10, 2026, a network of specialized brokers is systematically vacuuming up unspent credits issued to early-stage startups and offloading heavily discounted inference off-market. What began as casual voucher-swapping in backchannel founder chats has scaled into a multi-tiered shadow industry spanning direct sales outreach, specialized digital storefronts, and pseudo-enterprise proxy routers.

Founders regularly field unsolicited pitches promising 40 to 50 percent off standard list prices for off-market capacity from providers like Anthropic. During direct outreach to test the supply chain, Lenhard confirmed that individual brokers claim throughput capabilities of up to $100,000 in daily compute volume.

Rather than transferring master credentials or handing raw API keys to downstream buyers, these vendors operate through proxy relays. Their middleware intercepts inbound client requests, dynamically rotates across a pool of acquired developer keys, and routes the inference traffic directly to upstream providers.

Intermediaries and Infrastructure Risks

Public storefronts have emerged to formalize this pipeline across major model providers. Platforms such as AI Credits and AICreditMart position themselves openly as secondary compute exchanges, advertising discount listings between 30 and 80 percent alongside structured seller-onboarding workflows.

Alongside public exchanges, another tier of resellers dresses up its operation as legitimate enterprise aggregation. Outfits like CheapCredits, Tokvana, and Neokens market themselves as high-volume intelligent routers. CheapCredits, for example, lists a flat 40 percent discount across its entire model catalog. As Lenhard notes, securing a legitimate 40 percent volume discount from foundational providers is virtually unheard of unless an enterprise commands nine-figure run-rates.

For engineering leadership and security teams, funneling production inference through these unverified proxy layers introduces acute architectural vulnerabilities. Placing an unvetted intermediary in the path of enterprise traffic establishes a textbook man-in-the-middle vector. The broker's relay maintains full technical capability to intercept, log, inspect, and retain raw prompt payloads, proprietary system prompts, and structured model outputs prior to forwarding.

Beyond data exposure, the underlying foundation violates foundational model terms of service. Providers routinely purge arbitrated account pools without notice, leaving downstream clients with immediate outages, zero uptime recourse, and voided compliance postures. Chasing bottom-tier inference rates through arbitrated startup vouchers turns enterprise security into collateral damage just as model vendors prepare to clamp down on grant allocations.

Large Language ModelsCybersecurityAI in BusinessAnthropicCloud Computing