Jump to a section
AI field report · September 28, 2026 · Releases, economics and practical use
September’s AI news is about a wider set of choices. New hosted models, downloadable weights, smaller local models and personal agents are arriving together. The useful question is how much reliable work each can deliver, under what conditions, and at what total cost.
The catalyst for this review is All-In’s September 26 episode on model releases, pricing and frontier-company economics. Its strongest theme is that cheaper capability can move competitive advantage toward the software, distribution and controls surrounding a model. This article tests that argument against release announcements, model cards, licenses, pricing documentation and operational evidence.
1. What actually changed?
Three developments matter together. First, more capable models are available across a wider price range. Second, access is diversifying: a developer can buy inference through an API, operate some models on rented servers, or run a sufficiently small model on a suitable device. Third, the product increasingly includes tools, persistent state, permissions and a user interface that lets the model complete work.
These developments create pressure on undifferentiated model access. They also create new demands. A company still needs to connect data, define acceptable actions, evaluate results, monitor costs and support users. A lower token bill can make a workflow economical, but it does not automatically make the workflow dependable.
That framework also prevents a common category error. A foundation model, a coding assistant, an image generator and a personal agent can all be described as “AI releases,” yet their outputs, risk profiles and billing units differ. The comparison below starts by identifying what each release actually is.
2. A corrected map of the new releases
The episode’s rapid roundup compresses several launch dates and product families. The corrected board separates the model from the application and names the actual downloadable variant. It is a selected map of releases relevant to the episode, not an exhaustive census of every AI product.
11 releases shown
Sep 2 · Language model
Muse Spark 1.3
Hosted API
Hosted model; planned weights are not yet a released checkpoint.
Proprietary hosted service; open weights planned
Muse Spark 1.3 sourceSep 8 · Agent product
Muse personal agent
Application powered by Muse Spark
Tools, memory and service connections around a hosted model.
App launch is separate from the September 2 model release.
Muse product sourceSep 10 · Language model
DeepSeek V4.1 Flash
MIT weights + hosted API
Large MoE model; peak and off-peak API tariffs differ.
MIT
DeepSeek V4.1 Flash sourceSep 17 · Language model
Ternary Bonsai 2 27B
Local weights
Compact packed weights; total runtime memory is larger.
Apache 2.0
Ternary Bonsai 2 27B sourceSep 20 · Image model
Qwen-Image-2.1
Research-licensed weights
Image generation and editing; commercial use needs a separate license.
Qwen Research License; non-commercial only without separate license
Qwen-Image-2.1 sourceSep 21 · Language model
Grok 4.7
Hosted API
Text output with image input; long-context pricing changes above 200K.
Proprietary hosted service
Grok 4.7 sourceSep 22 · Language model
MiMo V2.6 Pro
MIT weights + hosted API
1.02T total / 42B active parameters; native multimodal input.
MIT
MiMo V2.6 Pro sourceSep 22 · Language model
MiMo V2.6 Flash
MIT weights + hosted API
309B total / 15B active parameters; distinct from Pro.
MIT
MiMo V2.6 Flash sourceSep 22 · Language model
Claude Opus 5.5
Hosted API
Adaptive thinking; workload savings differ from tariff reductions.
Proprietary hosted service
Claude Opus 5.5 sourceSep 22 · Language model
GPT-6 Sol
Hosted API
Higher capability tier than Luna; long-context premium applies.
Proprietary hosted service
GPT-6 Sol sourceSep 22 · Language model
GPT-6 Luna
Hosted API
Lower-cost family tier; native text output, image tools separate.
Proprietary hosted service
GPT-6 Luna sourceUse the filters to narrow the board. Dates describe the identified announcement or documented rollout; regional and account availability may differ. Follow the primary-source links for current conditions. No card is a recommendation or a claim that one model wins every task.
3. What the open-weight releases do—and do not—change
DeepSeek: the serving system matters alongside the model
DeepSeek V4.1 Flash was announced September 10. Its official release describes a 552-billion-parameter system with 8 billion active parameters for input processing and 16 billion for output, a one-million-token context window and visual understanding. The weights carry an MIT license. DeepSeek release; model card.
The company emphasizes its asymmetric architecture and reduced context-cache requirements. These are serving-efficiency claims, not evidence that the entire model fits on a typical laptop or that aggregate GPU demand must decline. The published API output tariff is $1.20 per million tokens at peak rates and $0.60 off-peak. DeepSeek pricing.
A lower inference price can support more retrieval passes, validation steps or background processing within the same budget. Yet those extra steps consume time and resources. The practical comparison should hold the task constant: identical input material, required output, tool access and quality threshold. Otherwise a cheap short answer is being compared with an expensive completed workflow.
Memory efficiency can change how much concurrent work a server can accommodate. It is still only one component of deployment cost. Weights, runtime buffers, long-context state, batching and latency constraints all matter. An architecture improvement cannot be converted into a universal hardware-cost reduction without specifying the serving setup.
MiMo: distinguish total parameters from active parameters
Xiaomi released MiMo V2.6 Pro and Flash on September 22. Both support text, images, video and audio, offer a one-million-token context window, and have MIT-licensed weights. Pro has 1.02 trillion total parameters and 42 billion active; Flash has 309 billion total and 15 billion active. The 309-billion figure discussed in the video belongs to Flash. MiMo announcement; Pro card; Flash card.
The benchmark evidence is more nuanced than a single headline. Xiaomi’s Pro card reports 89.9 on Terminal-Bench 2.1 against 89.1 for Opus 5, but 34.9 on Terminal-Bench 4.0 against 49.0 for that comparator. Those are vendor-reported results on different benchmark versions, not our tests. They illustrate why “matches a frontier model” needs a named task, version and setup.
In a mixture-of-experts system, a subset of the model’s parameters is active for a token. That can reduce the arithmetic used at each step, but the other weights still have to be stored and accessed appropriately. “Active parameters” is therefore not the model’s complete memory requirement. It also does not mean that two models with equal active counts have equal capability or cost.
The commercial value of downloadable weights is optionality. An operator can choose its hosting environment, customize parts of the system and reduce dependence on a single API endpoint. The tradeoff is operational responsibility: capacity planning, inference software, security, model updates and service reliability move toward the operator.
Bonsai: a smaller artifact changes the local-use discussion
PrismML’s September 17 Ternary Bonsai 2 27B release targets local deployment. Its Apache-licensed model is derived from Qwen3.8-27B. The detailed card lists a 5.95 GB packed language-weight file, a separate 7.21 GB packing option and an optional 0.63 GB vision component. Runtime memory and the context cache are additional. Bonsai announcement; Bonsai formats and model card.
The vendor’s 98.2% performance-retention claim summarizes 14 thinking-mode benchmarks against its own FP16 baseline. It does not mean 98.2% of every frontier model’s capability. The comparison is useful evidence about compression under the reported conditions; broad capability, latency and device suitability still require task-specific evaluation.
A compact model can be useful for document triage, drafting, extraction or offline experimentation when it satisfies the task. Local execution can reduce network dependence and limit transmission to a hosted model. Privacy still depends on the surrounding application, its logging, connected tools and any network access; the location of model weights alone is not a complete privacy guarantee.
For a local deployment, measure the whole experience: memory available while other applications run, time to first output, sustained generation speed, long-document behavior, battery or power use, and recovery from errors. A model that technically loads is not necessarily comfortable to use throughout a workday.
4. Image models belong in a different comparison
Qwen-Image-2.1 arrived September 20 with generation and editing in one system, native RGBA output and support for up to ten reference images. The advertised seven-billion-parameter figure describes its visual generation component, not a guaranteed complete runtime footprint. Qwen announcement; model card.
The license materially changes the open-source framing in the episode. The downloadable weights use the Qwen Research License for non-commercial research and evaluation. Commercial use requires a separate license. No independent like-for-like comparison establishing universal superiority over other image models was verified. Read the actual Qwen license.
For an image model, evaluate prompt adherence, text rendering, edit consistency, identity preservation, resolution and latency on the same brief. A compelling showcase can demonstrate possibility without establishing repeatability. Count retries and manual cleanup, especially when an output will be published or used in a customer workflow.
Rights are a separate procurement question. A downloadable checkpoint does not automatically provide unrestricted commercial permission. The relevant model license, any separate commercial agreement and the rights to input assets should be checked together. Hosting an image model yourself changes where inference occurs; it does not erase its license.
There is also no sensible single token-price leaderboard that puts image generation beside text reasoning. Image cost depends on the provider’s billing unit, resolution, quality setting, edit path and number of acceptable results. A useful production metric is cost per approved image at the required specification.
5. The hosted frontier is responding too
Claude Opus 5.5, September 22: Anthropic reports a one-million-token context window, up to 128,000 output tokens and always-on adaptive thinking. Standard input/output rates are $4/$20 per million tokens. Its claim of approximately 40% lower typical workload cost combines pricing with token efficiency; the list input/output rate reduction alone is 20%. Sonnet 5.5 and Haiku 5.5 were described as future releases. Opus announcement; Opus specifications and pricing.
GPT-6 Sol and Luna, September 22: OpenAI extends the family with lower-cost models using methods developed for Astra. Sol is $2/$10 per million input/output tokens; Luna is $0.10/$0.50. Both support approximately 1.05 million context tokens and 128,000 output tokens. Their native output is text; invoking an image tool is a separate capability. Sol and Luna announcement; dated rollout notice; Sol documentation; Luna documentation.
Grok 4.7, September 21: xAI’s text-output model accepts text and images and offers a 500,000-token context window. The base API rates are $2 input and $6 output per million tokens. Requests exceeding 200,000 input tokens have higher rates. The fast variant’s restricted distribution is different from the standard public API; speed claims should retain that distinction. Grok announcement; Grok documentation; developer release notes.
Tiering lets a provider serve both demanding and routine work. A user may choose an expensive model for an ambiguous research problem, a middle tier for most implementation work and a smaller model for repetitive classification. The important benefit is the ability to allocate capability where it changes the outcome.
Model selection is only one variable. Reasoning effort, tool budgets, context compaction, caching and the surrounding agent software can change quality, latency and the final bill. A nominally lower-priced model may generate more intermediate tokens or require more attempts. A higher-priced model may be economical when a single mistake creates substantial rework.
For a fair evaluation, keep a stable set of representative tasks and record both successes and failures. Include awkward inputs, incomplete instructions and long workflows. Measure the output a user actually needs, not just a benchmark score or a clean demonstration. Then compare complete cost and elapsed time at that acceptance threshold.
6. Meta’s Muse: separate the model from the agent
Muse Spark 1.3 is the model; Muse is the personal-agent product. Meta announced Spark 1.3 on September 2 and the agent app on September 8. “Muse Pop” in the episode title refers to momentum, not a model named Pop. Spark 1.3 is hosted through Meta’s API and Muse Code. Its planned open-weight release should not be described as already shipped. Spark model announcement; Muse app announcement.
Meta reports that its engineers used roughly 20% fewer tool calls and 25% fewer tokens with Spark 1.3 than with 1.2. That internal comparison concerns task consumption, not a same-sized reduction in the published token tariff. Separately, Meta’s August Muse Glimmer release provides 30-billion-parameter Apache-licensed weights for local agents; it should not be confused with hosted Spark. Muse Glimmer background.
The September 23 Connect announcements broadened the product story with voice, connectors, an agent email address and planned glasses access. Announced future features should not all be counted as delivered. Spotify independently confirmed a Muse connection for playback, saved music, playlists and discovery. This is evidence of integration and distribution, not a model benchmark. Meta’s September 24 recap; Spotify’s integration announcement.
The distinction explains why the personal-agent story can matter even when the underlying model is not the newest release. A useful application may reduce setup, connect existing services and turn a vague objective into a manageable sequence. Its value comes partly from how much coordination work it removes for the user.
That makes distribution and trust strategic variables. An agent that can reach email, calendars, documents and commerce services may save more time than a stronger isolated chatbot. It also has more opportunities to act incorrectly. Permission boundaries, previews, confirmations, audit trails and recovery paths become part of product quality.
Consider a travel task. Finding three suitable flights is an information problem. Booking one introduces dates, passenger details, payment, cancellation terms and authorization. A product that succeeds at the first task has not necessarily demonstrated safe competence at the second. Evaluate the entire sequence, including the point at which it asks the user to decide.
The podcast’s enthusiasm for mainstream adoption is a hypothesis worth watching. Durable evidence would be repeat use, completed tasks, retention and acceptable support costs. Download counts, anecdotes or a stock move can attract attention without proving those outcomes.
7. Read the price sheet before drawing a conclusion
A quoted API price usually applies to a particular model, token category and service condition. Uncached input, cached input and output can carry different prices. Long-context requests, off-peak windows, batch processing and data-contribution programs can change the bill further. A comparison that omits those conditions can be numerically correct and economically misleading.
Swipe or scroll the table horizontally to see all columns.
| Model | Input | Cached input | Output | Example bill* | Conditions |
|---|---|---|---|---|---|
| GPT-6 Luna | $0.1 | $0.01 | $0.5 | $0.0150 | Above 272K input tokens: 2x input/cache and 1.5x output for entire request |
| MiMo V2.6 Flash | $0.14 | $0.0028 | $0.28 | $0.0168 | Overseas standard API |
| DeepSeek V4.1 Flash | $0.3 | $0.006 | $1.2 | $0.0420 | Peak rates; off-peak rates are 50% lower |
| MiMo V2.6 Pro | $0.435 | $0.0036 | $0.87 | $0.0522 | Overseas standard API; not UltraSpeed |
| Muse Spark 1.3 | $1.25 | $0.15 | $4.25 | $0.1675 | Standard service |
| Grok 4.7 | $2 | $0.5 | $6 | $0.2600 | Base rates through 200K prompt tokens; above that 2x input/cache/output |
| GPT-6 Sol | $2 | $0.2 | $10 | $0.3000 | Above 272K input tokens: 2x input/cache and 1.5x output for entire request |
| Claude Opus 5.5 | $4 | $0.2 | $20 | $0.6000 | Standard service |
*Example bill: 100,000 uncached input tokens plus 10,000 output tokens. Arithmetic only, with no tools, retries, cache writes or long-context premium; it is not measured cost to complete a task. Base rates above are for standard synchronous requests at or below 200,000 input tokens. MiMo uses overseas rates, not domestic RMB or UltraSpeed.
Keep conditional offers separate: Muse Spark’s Contributor tier is $0.10 input, $0.002 cached input and $0.20 output, with permission for Meta to train on prompts and completions. Standard service has different data-use terms. DeepSeek off-peak rates are half its peak rates. Grok doubles all three token rates above 200K input; Sol/Luna double input/cache and multiply output by 1.5 above 272K input. Those premiums apply to the qualifying request, not just excess tokens.
Price references: MiMo · DeepSeek · Meta · xAI · Sol · Luna · Opus.
These are selected published rates, not matched-quality results. Prices do not measure reasoning quality, speed or reliability. A discounted data-contribution tier also changes the data-use agreement, so it should not be treated as the same product at a lower price.
Procurement should ask how the workload is billed in practice. Are reasoning tokens included in output? How often can input be cached? Are there request, concurrency or rate limits? Does an agent call paid tools? Are prices promotional or tied to a usage window? Those details matter more than a headline “half-price” claim without a specified predecessor and workload.
8. The useful denominator is a successful task
Suppose a cheap model produces an answer for one cent, but many attempts need repair. A more expensive model can be cheaper per accepted result if it needs fewer attempts or less human review. The reverse can also be true: routing easy tasks to a premium model may purchase capability the workflow never uses.
The calculator makes that tradeoff explicit. Its starting numbers are an illustrative scenario, not measurements of any named model. Success rate means the fraction of attempted jobs that meet a defined acceptance test after their allowed model calls. Average calls per job already includes retries. Human review is counted once per attempted job, avoiding a second multiplication by retries.
This deliberately simple model leaves out fixed platform fees, hosting, idle capacity, storage, retrieval, tools, taxes and the value of elapsed time. Add those costs before making a deployment decision. It also treats every successful job as equally valuable; real workflows need to account for the severity of different failure types.
A sensible pilot might use a low-cost model for extraction, escalate uncertain cases to a stronger model and require review before a consequential action. The resulting system should be tested as a system. Combining several individually good components does not guarantee that their failures will be independent or easy to detect.
9. Open-weight adoption is rising in a measured sample
Vercel’s September Production Index reports on its AI Gateway traffic through August. Open-weight models accounted for 56% of token volume and 14% of estimated spend. Average estimated price per token fell 23.2%; the median qualifying team’s decline was 7.6%. Vercel’s report and methodology.
One gateway · August 2026 · Share of each measure
Each full bar represents 100% of its own gateway measure. Spend uses published list prices; actual bills may differ.
The sample covers gateway traffic, not global market share. Tokens include input, output, reasoning and cache categories. The report uses a broader open-weight classification than earlier editions; its estimated spending is not audited revenue.
The strategic inference is narrower and more useful: cheaper models can capture substantial workload volume while premium models retain more spending per token. Volume leadership and economic value can diverge. For investors, neither statistic alone establishes profitability, customer loyalty or long-term pricing power.
10. What this means for AI business models
Model providers: price pressure increases the importance of serving efficiency, product differentiation and recurring demand. A provider can retain a customer while that customer moves to a cheaper tier, but revenue per task may still decline. Whether gross profit rises depends on cost reductions and additional useful consumption.
Application companies: cheaper capability can improve margins or enable features that were previously too expensive. The vulnerable business is one whose only value is forwarding a prompt. More defensible value may come from proprietary workflow context, integrations, trusted operation, distribution or a result that a customer can verify and pay for.
Cloud and infrastructure suppliers: efficiency changes demand rather than determining its direction. Lower cost can reduce compute required for a given task and make enough new tasks viable to increase total consumption. The outcome depends on adoption, utilization and willingness to pay. Neither “models are cheap, so data centers are unnecessary” nor “usage must grow without limit” follows from the release list.
Enterprises: model diversity improves bargaining options, but switching is not frictionless. Prompt behavior, tool interfaces, security reviews, evaluation sets, latency and data handling must be retested. The cost of switching may be small for a drafting tool and much larger for a regulated or deeply integrated workflow.
A useful company research map follows the flow of cash. Model suppliers sell inference or subscriptions; cloud operators sell usable capacity; chip and networking vendors sell equipment; applications sell an outcome or access to a workflow. These businesses can benefit at different times, and a strong order book at one layer does not prove satisfactory returns at every other layer. Related background: GPU rental economics, AI connectivity and the data-center field guide.
11. IPO headlines and infrastructure promises need different evidence
Confirmed filing status: Anthropic announced a confidential draft S-1 submission on June 1; OpenAI announced its submission on June 8. These statements do not establish completed offerings, fixed listing dates or final valuations. Anthropic filing announcement; OpenAI filing announcement.
The Wall Street Journal reported that Anthropic was considering November timing, with advisers seeing value in waiting for third-quarter financials. The accessible opening and a contemporaneous account do not establish that public safety rhetoric caused the shift. WSJ report (limited accessible text); MT Newswires’ account.
The Information reported a proposed founder voting-control structure; final terms require the actual offering documents. Anthropic’s existing Long-Term Benefit Trust arrangements concern board-selection rights and should be distinguished from shareholder votes and economic ownership. Reported governance proposal (headline access); Published trust structure.
Oracle is an execution-risk example. Bloomberg reported contingent payment protection tied to Project Jupiter’s delivery; that does not establish cancellation or a collapse in demand. Oracle’s September 14 statement distinguishes campus construction permits from the adjacent microgrid’s air-permit proceeding. The executed lease and notice were not reviewed here. Account of Bloomberg’s reporting; Oracle permitting statement.
An IPO discussion should separate the ability to raise capital from the quality of the underlying economics. Customers can value a product while public investors question cash requirements, concentration or governance. Conversely, a delayed listing does not by itself prove that the operating business has failed.
Infrastructure introduces another timing gap. A capacity agreement, a financing commitment, installed equipment and accepted service are separate milestones. A project can have genuine demand and still face delays that damage cash conversion. Evaluate contract conditions, commissioning, customer acceptance and the timing of payments before treating an announced capacity total as current revenue.
For this part of the AI cycle, a practical dashboard tracks contracted demand, available power, commissioned capacity, billable utilization, customer concentration, capex, debt obligations and operating cash generation. Those measures test the business more directly than a debate about whether an entire industry is a bubble.
12. Alignment and reliability are operating questions
Dario Amodei’s “We Must Pace the Frontier” proposes external evaluation and coordination; it does not promise a complete stop to training or releases. Anthropic’s September 18 Accenture announcement explicitly says model releases will continue. Its company-funded evaluation arrangement is a commitment to a process, not proof that a completed independent audit has established safety. Access, publication rights and conflicts of interest remain meaningful evaluation questions. Pacing proposal; Evaluation partnership.
METR’s August investigation describes agent coordination during a specific evaluation incident and discloses limits in the available evidence. That supports containment and adversarial testing; it cannot establish how often ordinary users will experience similar failures. METR incident investigation.
Meta’s Muse safety description includes a separate Sentinel for outgoing actions, approval requirements for sensitive operations, credential protection and an audit trail. Its confidential virtual-machine design was a future commitment, not a feature to assume already delivered at launch. These are specific architectural claims to evaluate, not a guarantee of alignment. Meta’s agent safety architecture; Muse launch and audit trail.
For an everyday deployment, translate the safety discussion into observable controls. Which data can the agent read? Which tools can it invoke? Which actions require approval? Can it contact arbitrary destinations? What happens if a retrieved document contains instructions aimed at the agent? Can the system explain what changed and help restore the prior state?
These controls have an economic role. They can add friction and review time, but a serious unauthorized action can overwhelm the savings from lower inference prices. The right measure is reliable completion within the permitted scope. A benchmark that rewards task completion without charging for destructive shortcuts can miss the property a real customer cares about most.
Model-level improvements and system-level controls should reinforce each other. It is unwise to assume the model will always recognize a malicious instruction, just as it is incomplete to assume permissions alone solve every reasoning failure. Evaluation should include both the model’s decision and the surrounding software’s enforcement.
13. AI-assisted science: meaningful progress, carefully bounded
Anthropic’s September 23 research announcement describes Claude helping identify an uncharacterized enzyme-associated system called ART around a previously known reverse transcriptase. The reported computational effort was approximately 950 agents over 21 hours, consuming 210 million tokens, with scientist direction and human laboratory work. The system’s function remains unknown. Anthropic’s research announcement.
The result is an early scientific lead reported by the company, not an independently replicated therapy, a validated replacement for CRISPR or an approved product. Similar-looking repeat structures can motivate research without establishing similar biological utility. The next evidence is functional characterization and reproducible experiments.
The investment question is how such workflows change the cost and speed of reaching a validated finding. Scientific usefulness depends on experimental design, reproducibility, measurement quality and expert interpretation. A large number of agents or tokens is an input measure; the quality of the resulting evidence is the output.
That creates opportunities across scientific software, data infrastructure and laboratory operations, but it does not justify assigning immediate therapeutic or commercial value to every announced discovery. Distinguish a computational hypothesis, an experimental result, a useful scientific explanation and an approved product. Each requires a different kind of proof.
14. What to watch over the next month
| Question | Evidence to request | Interpretation to avoid |
|---|---|---|
| Are cheaper models replacing expensive ones? | Same-customer workload migration at comparable quality, with actual bills. | Token share is identical to revenue or profit share. |
| Are personal agents useful beyond the demo? | Repeat use, accepted tasks, permission behavior, failures and support costs. | Downloads establish retention or dependable autonomy. |
| Does local inference work for the intended audience? | Full memory footprint, device latency, quality and operational burden. | A small weights file is the total system requirement. |
| Do frontier models preserve a premium? | Hard-task completion, fewer retries and measurable reduction in rework. | A single aggregate benchmark settles every workload. |
| Can infrastructure spending convert to cash? | Commissioned capacity, utilization, receipts and financing obligations. | Announced contracts are already earned revenue. |
| Does a release permit the planned use? | Current license, deployment terms and data-use conditions. | Downloadable means unrestricted commercial use. |
For a team selecting models, the next step is a bounded evaluation on its own work: define success, choose representative tasks, include failure cases, compare full costs and document permission boundaries. Repeat the evaluation when the model or agent software changes. A release cadence this fast makes a permanent winner list less useful than a repeatable decision process.
My interpretation is that value is becoming more closely tied to reliable execution. Cheaper intelligence broadens the set of economical tasks. The durable opportunity belongs to systems that turn that capability into work people trust, at a cost they can justify.
15. Sources and reading trail
The episode’s exported auto-generated transcript, description and chapter links informed this review. Names, dates and numerical claims were checked against the sources linked throughout this article. No direct product testing or independent model benchmark is claimed. Original graphics explain concepts; they do not depict measured model rankings.
The episode’s main sections are model performance and cost at 20:03, capital and competitive pressure at 29:22, personal agents and alignment at 1:07:58, and AI-assisted science at 1:27:13. These links preserve the source context; the analysis here is independently written.
- Model details: Follow the announcement, model-card, license and rate-sheet links in sections 2–7. A release date and a later page-update date are not interchangeable.
- Adoption: Vercel’s September Production Index measures its August gateway sample.
- Capital and governance: Company announcements establish confidential-submission status; the limited-access reporting in section 11 remains attributed reporting.
- Reliability: METR’s investigation and provider architecture disclosures inform the controls discussion.
- Science: Anthropic’s enzyme-system report is a company research account with unresolved biological function.