
openrouter.ai
August 20, 2026
3 min read
47/100
Summary
Ox Alpha is a free stealth reasoning model on OpenRouter for coding, sustained agentic work, production workloads, long-horizon software engineering, and complex reasoning. It accepts text, images, and video as input and returns text, supporting workflows that combine written and visual context. The model was released on August 20, 2026. An anonymous third-party provider develops and operates Ox Alpha during its preview period. OpenRouter routes requests directly to the model but is not its developer, owner, or provider. The provider retains prompts and completions, although it does not use them for training; other handling is governed by OpenRouter’s Stealth Model Terms. Ox Alpha is hosted by one provider, so OpenRouter has no alternative provider routing choices for requests. The model has a 1,048,576-token context window and can generate up to 131,072 completion tokens. It supports function calling through tools and tool_choice, plus JSON output through response_format without JSON Schema enforcement. OpenRouter lists prompt and completion token pricing at zero. Its reported median latency is 2.02 seconds and median throughput is 50 tokens per second. OpenRouter reports 99.99% uptime and 96.79% availability over the measured period.
Key Takeaways
What the discussion said
The thread treated Ox Alpha less as a model launch than as an opaque public beta: a free, anonymous endpoint whose real value may be the prompts and behavior it collects. Commenters broadly agreed that OpenRouter likely knows the provider while users do not, and that the arrangement gives a lab large-scale real-world testing without attaching early failures to its brand. Some saw that as a practical exchange for experimenting with an unreleased system; others argued that hidden provenance and unverifiable retention promises make it impossible to assess privacy, safety practices, or political constraints. Performance impressions were scattered but notable. One tester found its loose, creative reasoning unusually strong, even surpassing a respected competing model on tasks that had recently impressed them, while visual reasoning lagged. Others inferred a Chinese or GLM-family origin from its speed, terse output accounting, lengthy deliberation traces, and safety behavior, but those guesses remained speculative. Reports of refusals sharply conflicted: one person saw political censorship paired with permissiveness around harmful technical requests, while another encountered the reverse. Several readers stressed that external inference is perfectly defensible for public, low-stakes material, but the dominant mood warned against sending confidential work data to an unidentified provider. The thread also questioned why a provider would offer free access while retaining conversations if not to extract product intelligence.
Where opinion split
The central fight was whether an anonymous free model is an acceptable testing tool or an unacceptable data-risk black box. Defenders said public-data tasks carry little privacy exposure and real-world trials help providers harden models before release. Critics answered that users cannot verify retention, internal training, safety controls, or even the provider’s identity, so free access is a poor trade for anything sensitive.
Community Sentiment
Positives
Concerns