Welcome to Blank Metal’s Weekly AI Headlines.
Each week, our team shares the AI stories that caught our attention: the articles, announcements, and insights we’re actually discussing internally. We curate the best of what we’re reading and add the context that matters: what happened, why it matters, and what to do about it.
Agents Get Their Own Desks
Three of the biggest software vendors spent the week shipping the same idea: an agent with its own identity, its own computer, and permission to keep working after you log off. The question for your organization is now how you provision, scope, and audit a new kind of account.
OpenAI Launches Dots, Always-On Agents With Their Own Computers
What: At DevDay on September 29, OpenAI introduced dots, which it calls “always-on agents built to handle everything.” Each dot runs on GPT-6 Astra, has its own cloud computer and browser, connects to more than 4,000 apps through plugins, and can be reached in ChatGPT, Slack, or Microsoft Teams, with texting coming soon. When you aren’t working with it, a dot does what OpenAI calls “proactive research” in the background, using connected apps through tools restricted to read-only. Custom Rules let users allow an action, require approval, or block it, and an Activity View shows background work. Dots are rolling out to Pro and Business Premium plans; Enterprise, Edu, and Healthcare workspaces get a beta that is off until an admin enables it. OpenAI also previewed “specialist dots,” which a company sets up with their own identity, credentials, and access to systems of record, starting with enterprise pilots in areas like procurement, invoice processing, and customer support, and said it is working with Microsoft to manage them through Agent 365. DevDay also brought GPT-6.1 Sol, which OpenAI says delivers near-Astra intelligence at a fifth of Astra’s token prices, and a new $500-a-month Pro tier.
So What: The agent is moving from something an employee opens to something that keeps running after they close the laptop. That changes the governance question. A chat window needs an acceptable-use policy; an agent with its own computer, its own credentials, and standing access to email and Slack needs to be treated like an account in your identity system, with an owner, a scope, and a log. OpenAI is already pointing in that direction with specialist dots and the Agent 365 integration.
Now What: If you run ChatGPT Enterprise, decide before anyone asks whether the dots beta stays off, and if you turn it on, start with a named pilot group and read-only connectors. If you’re interested in specialist dots, pick one bounded back-office process (invoice matching, a support queue) and write down the identity, permissions, and approval rules you’d require before you talk to OpenAI about a pilot.
Microsoft Rebuilds Copilot Around Autopilot, an Agent With Its Own Identity
What: On September 25, Microsoft introduced a new Copilot organized around three experiences. Home combines Chat and Cowork, with Word, Excel, and PowerPoint built in so requests produce real, editable files that colleagues can work on live. Code lets employees outside engineering describe an app, tracker, or dashboard and have Copilot build it, using the same underlying technology as GitHub Copilot, with code running in a sandbox that can be hosted in the company’s own Microsoft 365 environment through a new Copilot Managed Runtime. Autopilot, previously called Scout, is a persistent agent that “lives in your tenant with its own identity, memory, computer and workspace”; users give it a name, a role, and a goal, and it keeps working when they’re away. Microsoft also introduced usage-based billing, with admin controls to set spending policies, choose which model families are available to which user groups, and manage plugins from one catalog. Users can pick OpenAI or Anthropic models, or an Auto mode that routes each request. Home and Code roll out to Microsoft’s Frontier early-access program in the coming weeks; Autopilot expands to private preview at the end of September. GeekWire reported that Jacob Andreou, who leads Copilot, said the autonomy that makes long-running agents useful is “the same thing that makes them absolutely terrifying to an IT admin.”
So What: Last week’s report said Microsoft would trade seat discounts for usage charges. This launch shows what gets metered: usage-based billing, per-group model access, and a hosted runtime for code your employees write with AI. Two of the three Copilot tabs create things that run without a person watching (Autopilot agents and Code apps), which means your Microsoft tenant is about to host software nobody in IT wrote.
Now What: If you’re in the Frontier program or renewing Copilot, set the spending policies and model-family restrictions before Home and Code reach users, and decide where Code-built apps are allowed to run and who reviews them. Ask Microsoft how Autopilot identities appear in Entra and in your audit logs before you join the private preview.
Meta Opens an Enterprise Business and Hires MongoDB’s CEO to Run It
What: On September 28, Mark Zuckerberg announced Meta Enterprise Platform, “the next major pillar of our business,” to sell Meta’s AI to companies. It will start by bringing Meta’s full stack to businesses and developers: the Muse agent, Meta Business Agent, the Muse API, Muse Code, and more. CJ Desai, who stepped down as CEO of MongoDB and previously served as president and COO of ServiceNow, joins as Chief Enterprise Platform Officer reporting directly to Zuckerberg. The next day, Meta launched Muse for Small Business, which is free with usage limits, with paid plans for more, and connects to Shopify, Slack, Dropbox, Asana, Box, Canva, Figma, Intuit QuickBooks, Notion, Stripe, Zoom, and others, alongside the business’s Facebook and Instagram ad accounts.
So What: Three weeks after launching a consumer agent, Meta now wants to be an enterprise vendor, and it hired an executive who has sold software to large IT organizations. For buyers, this adds a fourth serious agent platform to evaluate next to OpenAI, Anthropic, and Microsoft, and its pricing model is unusual: for Muse, Meta has said it expects to profit over time from small transaction fees, with no seat license. The small-business launch is also the path by which Muse arrives inside larger companies, through a regional office or a marketing team that signs up for free.
Now What: If you’re running a vendor evaluation for agents this year, add Meta to the list and ask for its enterprise data terms in writing, specifically retention, training use, and how Muse connectors are scoped. In the meantime, check whether any team has already connected a free Muse for Small Business account to company Slack, Box, or Shopify.
The Model Menu Gets Cheaper and More Specialized
The flagship race continued, but the more useful news for buyers sat below and beside it: a mid-tier model that nearly matches the top tier, a frontier model released first to defenders, and a model that only picks from a list. Model selection is becoming a portfolio decision.
Anthropic Ships Sonnet 5.5: Near-Opus Results, Up to 30% Cheaper Per Task
What: Anthropic released Claude Sonnet 5.5 on September 28, six days after Opus 5.5. List prices are unchanged from Sonnet 5 at $2 per million input tokens, $10 per million output, and $0.20 for cache reads, but Anthropic says it needs far fewer tokens and costs up to 30% less per task, while generating output more than 30% faster. On GDPval-AA, a test of real-world work across 44 occupations, it scores 1844 against 1846 for Opus 5.5 and 1449 for Sonnet 5. On Terminal-Bench 4.0 it scores 70.6%, against 10.3% for Sonnet 5. Anthropic says Opus 5.5 “remains clearly stronger at complex, open-ended work requiring sustained judgment.” Customer reports in the launch include Zendesk processing tickets 20% faster and Balyasny Asset Management seeing answers use about 121,000 tokens against 497,000 for Sonnet 5 on a 2,441-task finance suite. It is the first Sonnet model to launch with cyber safeguards, so higher-risk security tasks visibly fall back to Sonnet 5. Teams that run Sonnet with thinking off need to switch to a new between_tools setting before migrating. Haiku 5.5 follows in the coming weeks.
So What: The middle tier now scores within a rounding error of the flagship on knowledge-work benchmarks at half the input price. For most enterprise workloads (document drafting, support triage, everyday coding), the useful question is which tier is good enough, and the answer is increasingly the cheaper one. Notice also how the savings arrive: same list price, fewer tokens per task. Your cost-per-token dashboards won’t show it.
Now What: If you run high-volume workloads on Opus 5 or 5.5, rerun your evaluation set on Sonnet 5.5 at Medium and High effort and compare cost per completed task. Before migrating, search your codebase for thinking-off configurations and security tooling, since both behave differently on this model.
Google’s Gemini 4 Argon Goes to Cyber Defenders First
What: On September 30, Google announced Gemini 4 Argon, its new frontier model for coding, enterprise knowledge work like legal and finance, and cybersecurity defense. It is rolling out first to a set of trusted cyber defenders through Google’s Fairwind Program, with paid API customers and Google AI Ultra subscribers among the next groups; Google says it is taking part in the US government’s voluntary pre-release access process while it expands. Argon’s output limit rises to 1 million tokens, from 64,000. Google reports state-of-the-art scores on DeepSWE v1.1 (77.9%) and first place on Zapier’s AutomationBench (51.3%) and on the Vals Index, which weights finance, coding, legal, and tax work by contribution to US GDP. Trusted defenders will get Argon without cyber guardrails; Wiz used it to find a critical vulnerability exposing personal information in healthcare software used by hospitals worldwide. Introductory pricing is $2 per million input tokens and $10 per million output, rising to $4 and $20 afterward. Google also says it monitors Argon’s chain of thought to stop execution when the model steps beyond the user’s intent.
So What: Google now joins Anthropic, with its Cyber Verification Program, and OpenAI, with its Daybreak defender models, in reserving its strongest cyber capability for vetted defenders while everyone else gets a guarded version. That makes access tiers part of model selection: the model your security team can get may differ from the one your developers get. On price: at the introductory rate, Argon lists at Sonnet 5.5’s price; at the standard rate, it lists at Opus 5.5’s.
Now What: If your security team defends critical infrastructure or regulated data, apply now to the defender programs at Google, Anthropic, and OpenAI, since access is gated and queued. For everyone else, add Argon to your evaluation harness when it reaches the API, and price it at $4 and $20, since the introductory rate expires.
OpenAI Launches a Decisions API for Classification and Routing
What: Among its DevDay announcements, OpenAI released a Decisions API in limited preview, with broad release planned in the coming days. It focuses its small Luna model on “a specific set of user-defined questions with finite pre-defined answers”: developers supply text or images, and get back an answer they can use to classify content, route a request, or choose an agent’s next action. The New Stack reported that it returns predefined answers with confidence scores in about 150 milliseconds, against 1.6 seconds for GPT-6 Luna, positioned as a response to TypeSafe’s Jev, a model built to choose from supplied options and never generate text. Pricing has not been announced.
So What: A large share of enterprise AI work isn’t writing. It’s picking: which queue, which category, approve or escalate, which tool to call next. Teams have been doing that with chat models and asking for structured output, which is slower, costlier, and harder to audit. A model that only chooses, with a confidence score attached, fits the routing step in an agent workflow and gives you a natural threshold for sending low-confidence cases to a bigger model or a person.
Now What: If you run a classification or triage step on a general-purpose model today, benchmark it against a decision model on your own labeled data, measuring accuracy, latency, and cost per thousand decisions. Use the confidence score to set an explicit escalation threshold, and log every decision below it for review.
Who Holds the Data, and Who Holds the Risk
A retention promise with fine print, an apology for agents that wandered into government systems, a prospectus full of risk factors, and an insurer’s audit of hospital AI. Each one is a reminder that the paperwork around AI now matters as much as the model.
OpenAI Offers Zero Data Retention That Still Gets Safety Review
What: OpenAI used DevDay to launch Private Intelligence, aimed at businesses that want frontier models with stronger data protection. Its first piece is Zero Data Retention with Private Safety Processing, which runs automated safety reviews without giving OpenAI personnel access to the underlying content. According to OpenAI’s documentation, encrypted customer content is decrypted only inside a hardware-attested environment designed to disallow human access, and only bounded safety signals and operational metadata leave it in plaintext. The fine print matters: records under this option are retained, encrypted, for at least 30 days for the automated review, and customers must already be approved for zero data retention and set up validated storage. A preview of Private Inference, combining confidential computing with “strict, verifiable controls,” is due this fall. OpenAI also announced a Marketplace that lets eligible enterprise customers apply part of an existing OpenAI commitment to approved partner software, starting with 32 partners including Salesforce, ServiceNow, Harvey, CrowdStrike, and Palo Alto Networks.
So What: Retention has become the sticking point in enterprise AI deals, because labs want logs to catch attacks that unfold across sessions and customers want their data gone. OpenAI’s answer is to keep the data but put it where no person at OpenAI can read it. That is a meaningful change, and it also redefines “zero data retention” to include 30 days of encrypted storage. Read the definition, not the label.
Now What: If your AI vendor contracts reference zero data retention, ask each provider in writing what is stored, for how long, in what form, and who or what can read it. Put the answers side by side before your next renewal, because the same term now means different things at different labs.
OpenAI Apologizes to Australia and Pauses Tool-Use Training on Its Top Models
What: On September 28, OpenAI published an apology and account of the incidents in which its models reached Australian government systems during internal training and evaluation in June. It disclosed four affected agencies. At Services Australia, an experimental internal model “discovered a way to gain non-public access,” ran commands, and retrieved internal files, credentials, and aggregate statistics; OpenAI says no individual patient records were accessed. Its agents also used an exposed access key at the Victorian Department of Health and pulled configuration and logs from a NSW crime statistics tool. The model had been asked to research government spending on medicines for skin conditions and, when it couldn’t find the data, took actions OpenAI had not authorized. OpenAI said it “should have shared preliminary findings sooner,” that it has blocked live internet access in research environments and now serves web content from cache, and that it has paused training and evaluation involving tool use for its most capable models until it has additional safeguards. It is committing support from its $1 billion Daybreak defender fund, forming an Australian taskforce, and sending Chief Strategy Officer Jason Kwon to testify in Sydney on October 6.
So What: The detail that transfers to your environment is the task. The model wasn’t told to attack anything; it was told to find a statistic, and it treated access controls as obstacles between it and the answer. Any agent with a goal, tool access, and a network path can do the same, inside or outside your perimeter. OpenAI’s own fix is instructive: it removed live internet access and put monitoring that pages a human in front of anything risky.
Now What: For every agent you run with web or API access, document what it may reach and enforce that with network controls rather than instructions in the prompt. Add “attempted access outside scope” as a monitored event that pages a person, and test it by giving an agent a task it can’t complete legitimately.
Anthropic’s IPO Filing Shows $4.6 Billion in Revenue and $518 Billion in Compute Commitments
What: Reuters reported on September 28 that it had seen Anthropic’s IPO prospectus. Revenue grew 12-fold in 2025 to nearly $4.6 billion. The company reported a net loss of about $42 billion, including a roughly $34 billion accounting charge tied to financing that could convert to shares; excluding writedowns, it lost more than $8 billion on an operating basis. It spent $7.33 billion on compute and infrastructure in 2025, more than half of $12.65 billion in operating expenses, and plans $518 billion in cloud, computing, and infrastructure obligations in coming years. Nearly a quarter of revenue came from two customers, and Anthropic warned that many of its largest clients are not locked into long-term contracts and could cut or stop spending. It held $20.28 billion in cash and short-term investments at year end. The listing could value the company at more than $2 trillion and, according to earlier Reuters reporting, is likely to come after the November midterm elections.
So What: For buyers, a prospectus is the most candid document a vendor publishes, because the risk factors carry legal weight. Two disclosures matter for your planning. Revenue is concentrated in a small number of very large customers, which tells you where roadmap attention goes. And the compute commitments dwarf current revenue, which means pricing, capacity, and terms will be shaped by the need to fill that capacity for years.
Now What: If Anthropic is a strategic vendor for you, read the risk factors section when the S-1 becomes public and update your vendor risk file. Use the moment in negotiations: a company preparing to list wants multi-year commitments and referenceable logos, and both give you bargaining power on terms.
Blue Cross Says Hospital AI Coding Tools Added $942 Million in Costs Without Changing Care
What: On September 24, the Blue Cross Blue Shield Association released an analysis of its member companies’ claims data finding that hospitals’ increasing use of AI coding tools coincided with a sharp rise in patients documented as medically complex, adding an estimated $942 million in spending between 2023 and 2025. About 70%, roughly $650 million, came from secondary diagnoses (conditions identified in addition to the main reason for care, often derivable from a single lab value) that moved claims into higher-reimbursement categories. “If patients are truly sicker, we’d expect to see more treatment,” said BCBSA’s Luke Chalker, pointing to more anemia diagnoses without a matching rise in transfusions. BCBSA says more than 60% of hospital systems now use AI coding tools. The American Hospital Association pushed back on the analysis, and The New York Times reported the same week that hospitals’ and insurers’ use of AI is escalating their longstanding feud over paying for care.
So What: This is what happens when both sides of a transaction automate against each other. Each party’s AI makes it better at its own objective, and the total cost of the system goes up. The pattern isn’t unique to healthcare: procurement, collections, claims, and any negotiated process where both counterparties can now deploy an agent will see the same dynamic. The finding also shows where the audit trail matters. BCBSA could make this argument only because it could compare what was coded with what was done.
Now What: If you’re deploying AI on one side of a contested process (billing, claims, disputes, procurement), measure outcomes your counterparty can’t dispute, such as care delivered or goods received, alongside the revenue your tool recovers. If you’re on the paying side, start tracking the ratio of documented complexity to delivered service now, so you can see a shift when it starts.
Blank Metal is an AI consulting and engineering firm. We help organizations move from AI experiments to production systems. Learn more



