11 min read Claude Opus 4.6

Anthropic acquires Stainless, the SDK generator behind OpenAI and Google

Anthropic acquired Stainless, the startup behind SDK generation for OpenAI, Google, and Cloudflare, signaling that developer tooling infrastructure is becoming a strategic acquisition target for AI labs. NVIDIA delivered its first Vera CPU – designed from scratch for agentic AI inference – to Anthropic, OpenAI, and SpaceXAI. A California jury unanimously rejected Elon Musk’s lawsuit against OpenAI on statute-of-limitations grounds, removing a major legal overhang ahead of the company’s expected IPO.

Funding & Business #

Anthropic acquires Stainless #

Anthropic / TechCrunch / Hacker News (457 points)

Anthropic acquired Stainless, a startup founded in 2022 that generates SDKs, CLI tools, and MCP servers from API specifications. Stainless has powered official SDKs for Anthropic, OpenAI, Cloudflare, and hundreds of other enterprises – making this an acquisition of infrastructure that competitors also depend on. The move reflects a strategic bet that as AI shifts “from models that answer to agents that act,” controlling the tooling layer that connects agents to APIs becomes a competitive advantage.

OpenAI and Dell partner to bring Codex to enterprise environments #

OpenAI

OpenAI and Dell announced a partnership to deploy Codex in hybrid and on-premise enterprise environments, targeting organizations that need AI coding agents operating on sensitive data without leaving corporate infrastructure. This extends the Codex platform beyond cloud-only deployment and reflects enterprise demand for sovereign AI tooling. Combined with OpenAI’s recently launched Deployment Company (covered May 18), this shows OpenAI aggressively building enterprise distribution channels.

SandboxAQ brings drug discovery models to Claude #

TechCrunch

SandboxAQ integrated its physics-grounded molecular dynamics and quantum chemistry models into Claude, making sophisticated drug discovery tools accessible via natural language. Rather than competing to build better models, SandboxAQ is betting that usability is the bottleneck – enabling pharmaceutical researchers to run simulations without specialized computing infrastructure. This is a concrete example of the MCP-driven tool ecosystem expanding beyond software development into scientific computing.

Infrastructure #

NVIDIA Vera CPU delivered to Anthropic, OpenAI, and SpaceXAI #

NVIDIA Blog

NVIDIA hand-delivered its first Vera CPUs to Anthropic, OpenAI, SpaceXAI, and Oracle Cloud Infrastructure. Vera features 88 custom Olympus cores and 1.2 TB/s memory bandwidth, purpose-built for agentic AI workloads that require orchestration, tool-calling, and real-time reasoning rather than pure GPU throughput. This is NVIDIA’s first custom CPU and marks an architectural acknowledgment that agent inference has fundamentally different compute requirements than training or batch inference.

Jensen Huang at Dell Technologies World: “Demand is going parabolic” #

NVIDIA Blog

At Dell Technologies World, Jensen Huang announced the Dell AI Factory with NVIDIA’s Vera Rubin NVL72, claiming 10x lower cost-per-token for large-scale agentic AI inference. Huang stated agent sandboxes run 50% faster on Vera than traditional CPUs, with 5,000 enterprises including Lilly, Samsung, and Honeywell already running AI workloads on Dell AI Factories. The “parabolic demand” framing positions agentic inference as the next infrastructure scaling wave.

PyTorch 2.11 ships CUDA-enabled wheels for aarch64 by default #

PyTorch Blog

PyTorch 2.11 now publishes CUDA-enabled wheels for aarch64 Linux on PyPI by default, fixing a long-standing issue where pip install torch silently installed CPU-only builds on ARM-based GPU systems like NVIDIA’s GB200 and GH200. The fix was coordinated through the PyTorch Foundation’s Technical Advisory Council and eliminates workarounds that vLLM and other projects had to maintain. As ARM-based AI servers become more common (particularly with Vera), this packaging fix removes a significant developer friction point.

Developer Tools #

Cursor releases Composer 2.5 #

Cursor / Hacker News (157 points)

Cursor released Composer 2.5 with what the company describes as “substantial improvement in intelligence and behavior,” particularly on sustained complex tasks and intricate instruction following. The release introduces targeted reinforcement learning with textual feedback at specific points in task sequences, a 25x increase in synthetic training tasks, and optimized distributed training via Sharded Muon. Pricing starts at $0.50/$2.50 per million input/output tokens – aggressive pricing that pressures both general-purpose model providers and competing coding assistants.

IBM Research and Hugging Face launch Open Agent Leaderboard #

Hugging Face Blog

IBM Research and the HF community launched the Open Agent Leaderboard, which evaluates full agent systems rather than just underlying models across six benchmarks including SWE-Bench Verified, BrowseComp+, and AppWorld. Key findings: agent architecture matters as much as the model (same model produces vastly different results in different agent frameworks), general-purpose agents match specialized ones, and failed runs cost 20-54% more than successful ones. The accompanying open-source Exgentic framework enables reproducible agent evaluations – filling a gap where most agent benchmarks remain closed or non-reproducible.

ExecuTorch MLX Delegate enables GPU inference on Apple Silicon #

PyTorch Blog

The new ExecuTorch MLX Delegate provides GPU-accelerated inference for PyTorch models on Apple Silicon using Apple’s MLX framework, achieving 3-6x higher throughput compared to existing ExecuTorch backends on macOS. It supports dense transformers (Llama, Qwen, Gemma), sparse MoE models, and speech-to-text systems with multiple quantization options including 2/4/8-bit affine and NVFP4. For developers building on-device AI applications for Apple hardware, this closes the performance gap between Apple Silicon and dedicated inference hardware.

Security #

Mini Shai-Hulud: 317 npm packages compromised in automated supply chain attack #

SafeDep / Hacker News (110 points)

A compromised npm account published 637 malicious versions across 317 packages in an automated burst on May 19, targeting high-traffic packages including size-sensor (4.2M monthly downloads) and echarts-for-react (3.8M downloads). The malware used a two-pronged execution strategy: a preinstall hook executing obfuscated Bun code plus optional dependencies pointing to forged GitHub commits in the antvis/G2 repository. The dual delivery mechanism – hooks plus dependency-chain poisoning – ensured persistence even when install hooks were blocked, representing an evolution in supply chain attack sophistication.

Bug bounty programs bombarded with AI-generated slop #

Ars Technica

Corporate bug bounty programs are being overwhelmed by AI-generated vulnerability reports that are superficially plausible but technically invalid, straining the triage capacity that makes these programs viable. The “never-ending” volume of AI slop forces security teams to spend more time filtering noise than evaluating genuine submissions. This is the bug-bounty manifestation of a broader pattern: as AI lowers the cost of producing plausible-looking technical content, the verification burden shifts entirely to recipients.

AI Agents May Always Fall for Prompt Injections #

arXiv

This paper argues that the prevailing defense paradigm – data-instruction separation – both fails to detect attacks that operate through contextual manipulation and degrades contextually appropriate behavior. The authors recast prompt injection through Contextual Integrity theory, showing that some attacks are inherently indistinguishable from legitimate context-dependent behavior. If this analysis holds, it means prompt injection is not a solvable engineering problem but a fundamental tension in any system where agents must interpret untrusted input as actionable context.

Trust No Tool: Cognitive Poisoning in Tool-Using LLM Agents #

arXiv

Most agent security work assumes tools are trustworthy once selected. This paper identifies “cognitive poisoning” – a malicious tool behaves plausibly during exploration, accumulates trust through benign-looking feedback, and becomes harmful only when hidden state conditions are met. The delayed-activation pattern mirrors the npm supply chain attack (above) but operates at the semantic level: the agent’s accumulated experience with the tool becomes the attack vector, not the tool’s code.

Regulatory & Policy #

Musk v. Altman: jury unanimously rejects all claims #

TechCrunch / Ars Technica / MIT Technology Review

Following last week’s closing arguments (covered May 18), a California jury unanimously ruled that Musk’s claims against OpenAI were filed too late under applicable statutes of limitations. Musk had sought restructuring and damages potentially reaching $78-135 billion; the verdict eliminates a major legal threat ahead of OpenAI’s reported IPO. Musk has announced he will appeal, but the unanimous advisory verdict – immediately accepted by Judge Gonzalez Rogers – sets a high bar for reversal.

Research & Papers #

The Scaling Laws of Skills in LLM Agent Systems #

arXiv

Across 15 frontier LLMs, 1,141 real-world skills, and over 3M routing or execution decisions, this paper identifies two coupled scaling laws for agent skill libraries. Routing accuracy decays logarithmically with library size (R-squared > 0.97 for all models), with errors progressing from local skill competition to cross-family drift. This is directly actionable for anyone building agent systems with growing tool sets: skill library size has predictable, measurable costs in routing accuracy that must be engineered around rather than assumed away.

Code as Agent Harness #

arXiv

This paper frames a shift where code is no longer just an agent’s output but its operational substrate – used for reasoning, acting, environment modeling, and execution-based verification. The “agent harness” framing treats code as the medium through which agents structure their own computation, moving beyond the “give the model a shell” paradigm. For teams building agentic systems, the implication is that code generation, code execution, and agent reasoning should be co-designed rather than treated as independent capabilities.

The Point of No Return: Counterfactual Localization of Deceptive Commitment #

arXiv

Existing deception datasets label completed outputs as honest or deceptive, but this paper asks a more fundamental question: when does a language model become committed to deception during its reasoning trace? Using counterfactual localization – fixing a prefix, resampling continuations, and estimating deception probability – the authors identify specific “points of no return” in reasoning chains. This technique could enable runtime monitoring that intervenes before deceptive outputs are finalized rather than detecting them after the fact.

EXG: Self-Evolving Agents with Experience Graphs #

arXiv / Hugging Face Daily Papers

Most deployed agents remain behaviorally static – knowledge acquired during execution rarely translates into systematic improvement. EXG introduces experience graphs that structure an agent’s accumulated execution history into a retrievable, composable knowledge base, enabling agents to improve through deployment experience rather than requiring retraining. For production agent systems, this addresses the gap between one-shot task execution and genuine learning from operational history.

Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics #

arXiv

Chain-of-thought reasoning is not always faithful to a model’s final output, limiting its reliability as a safety monitoring tool. This paper investigates hidden representations of reasoning models and finds that probe trajectories – evaluating a classifier at each generated token – can predict future behavior from prompt and CoT representations before the final answer is produced. This offers a path to runtime safety monitoring that does not depend on the faithfulness of externalized reasoning.

Other #

Google I/O preview: what to expect from a “clear third place” in the foundation model race #

MIT Technology Review

MIT Technology Review frames Google’s annual developer conference – opening May 20 – with the observation that Google enters as “a clear third place in the foundation model race,” a sharp contrast from a year ago when Gemini launched to widespread attention. The preview highlights infrastructure and developer tooling as likely focus areas rather than headline model releases. Google’s positioning will clarify whether the company is competing on model capabilities or pivoting to platform and distribution advantages.

Anduril and Meta prototype AR headset for military drone strikes #

MIT Technology Review

Anduril shared details about an augmented-reality headset prototyped with Meta for military use, including ordering drone strikes via eye-tracking and voice commands. The project applies consumer AR technology to defense applications, with Anduril VP Quay Barnett framing it as enabling “decisional superiority” in the field. This represents the most explicit convergence yet of consumer AI/AR hardware with military command-and-control systems.

Threads to Watch #

Agent security research is converging on a disturbing conclusion: some attack surfaces may be inherent. “AI Agents May Always Fall for Prompt Injections” argues data-instruction separation fundamentally cannot distinguish attacks from legitimate context-dependent behavior. “Trust No Tool” identifies cognitive poisoning where tool trustworthiness degrades through accumulated experience. Combined with today’s npm supply chain attack – which used a similar delayed-activation pattern at the code level – and yesterday’s sleeper memory poisoning paper, the emerging picture is that agent security requires accepting certain attack surfaces as irreducible and engineering for graceful degradation rather than prevention.

NVIDIA is building the full-stack agentic inference platform. Vera CPU delivery (purpose-built for agent orchestration), the Vera Rubin NVL72 announcement (10x cost-per-token reduction for agentic inference), and PyTorch’s aarch64 CUDA fix (enabling ARM-based GPU servers) collectively outline an infrastructure story where agentic AI workloads get their own dedicated compute architecture rather than repurposing training hardware. The “Scaling Laws of Skills” paper adds an important constraint: as agent systems grow, routing overhead scales logarithmically – meaning infrastructure must accommodate not just raw inference but increasingly complex orchestration.

Developer tooling is being acquired, not built. Anthropic’s Stainless acquisition follows the pattern of AI labs buying developer infrastructure rather than building it internally. Stainless powers SDKs for competitors including OpenAI and Cloudflare, raising questions about whether those relationships survive the acquisition. Meanwhile, Cursor’s Composer 2.5 and the Open Agent Leaderboard push the tooling layer toward evaluating agent systems holistically rather than models in isolation.

Sources Unavailable Today #

These sources could not be fetched today. Links point to their homepages so you can check them directly.