The first 500 users of Architect’s new LLM inference tool get $20 free credit. That’s one small perk for trying Liquid Inference, a product built by a trading firm—not an AI lab—that’s changing how LLM inference is priced.
What is Liquid Inference?
Liquid Inference is an exchange-style router for LLM inference. It comes from Architect Financial Technologies, which runs the AX perpetual futures exchange. In May 2026, the firm acquired a US Designated Contract Market to list GPU compute futures, though regulatory approval is still pending. The team built Liquid Inference using their experience building financial exchanges, creating a two-sided price discovery system for LLM inference. Unlike existing LLM API routers, it runs a real-time auction for every single prompt request.
How the per-request auction works
The process has four clear steps. First, a client sends a standard OpenAI or Anthropic API call, no full code rewrite needed—only swapping a base URL. Second, the buyer’s constraints filter out offers that don’t meet their rules, such as per-job cost caps, time to first token limits, required regions, zero data retention, or approved providers and models. Third, providers quoting the requested model compete in an auction; the lowest qualifying offer wins. Fourth, the max price is locked before generation starts, and billing only covers the actual metered usage of the model.
Buyers have extra controls too. They can use an Auto mode to pick the best model for a given task, plus preset routing rules and full multi-modal support, per a LinkedIn post from Harrison. Account holders get access to live order books, per-provider and per-model quotes, and cleared transaction records—this level of real market data is rare for LLM APIs.
Who benefits from this auction model?
For developers and LLM buyers, Liquid Inference is drop-in compatible with popular agentic coding tools: Claude Code, Codex, OpenCode, Cursor, Pi, and Cline. It also has a referral program: users get 20% of their referred fees as free inference, plus an extra 10% on second-level referrals.
For LLM inference providers, onboarding takes minutes, not weeks. All prompts use the OpenAI API standard, so integration is straightforward. Providers can adjust their quotes based on their own costs, letting them sell spare GPU capacity only when it makes sense for their business. Payouts run through Stripe, with itemized records for every job.
Compared to existing tools like OpenRouter and Hugging Face Inference, Liquid’s core difference is the per-request auction. OpenRouter uses price-weighted load balancing, while Hugging Face picks the fastest provider by default (cheapest is an optional suffix). Liquid says it has hundreds of models, though its full provider list isn’t public. OpenRouter has 500+ models and 80+ providers, Hugging Face has 18 listed partners. Liquid also stands out for its live public market data, which neither OpenRouter nor Hugging Face offers.
Still, some key details are missing. We don’t have full info on Liquid’s platform fees, latency data, or complete provider lineup—all of which would matter for developers choosing between tools.
The question now is: will LLM developers opt for this auction-based model over the simpler, more established load-balanced routers? For teams that prioritize getting the lowest possible price per request, or want granular control over inference rules, Liquid might fill a gap. But without full latency and fee data, it’s too early to say if it will win over broad adoption.
素材来源:MarkTechPost · AI情报、大模型、AI应用落地、AI政策与监管
查看报道原文

发表第一条评论吧