Anthropic secures SpaceX's Colossus 1 data center and its 220,000 GPUs
Anthropic secured SpaceX’s entire Colossus 1 data center – over 300 megawatts and 220,000 NVIDIA GPUs – while unveiling Managed Agents “dreaming” and autonomous outcome-setting at its Code with Claude developer event. SpaceX separately proposed a $119 billion “Terafab” chip factory in Texas with Intel, DeepSeek targets a $45 billion first venture round with Chinese state backing, and new research introduced real-time safety interception for agent tool use alongside saturation-resistant multi-agent benchmarks.
Infrastructure #
Higher usage limits for Claude and a compute deal with SpaceX #
Anthropic / Ars Technica
Anthropic secured access to SpaceX’s entire Colossus 1 data center in Memphis – over 300 megawatts and 220,000 NVIDIA GPUs – available within one month, immediately doubling Claude Code’s five-hour rate limits across Pro, Max, Team, and Enterprise plans and removing peak-hours reductions. The agreement also includes exploring “multiple gigawatts of orbital AI compute capacity,” adding to Anthropic’s existing deals with Amazon (5 GW), Google/Broadcom (5 GW launching 2027), and Microsoft/NVIDIA ($30B Azure capacity partnership). The scale of compute procurement – Anthropic alone now has line-of-sight on more than 10 GW of capacity – reflects how quickly inference demand is outpacing supply.
SpaceX may spend up to $119B on ‘Terafab’ chip factory in Texas #
TechCrunch
SpaceX filed a proposal in Grimes County, Texas for a “multi-phase, vertically integrated semiconductor manufacturing” facility with Intel, starting at $55 billion and potentially reaching $119 billion across phases. Musk argues semiconductor manufacturers aren’t producing chips fast enough for AI and robotics demands, with the factory intended to serve SpaceX satellites, Tesla autonomous vehicles, and space-based data centers. The vertical integration play – an AI-adjacent company building its own chip fab – would represent one of the largest single manufacturing investments in history, though Musk’s semiconductor manufacturing track record is untested.
TSMC taps wind power as AI chip demand soars, Taiwan feels energy crunch #
Ars Technica
TSMC is procuring wind power as record AI chip manufacturing demand strains Taiwan’s electrical grid, highlighting a constraint that chip makers face globally. AI workload growth requires not just more fab capacity but more power, and grid limitations are becoming a binding constraint on semiconductor production timelines.
AI boom pushes Samsung to $1T #
TechCrunch
Samsung crossed the $1 trillion valuation mark on AI-driven chip demand, becoming only the second Asian company after TSMC to reach the milestone. The valuation concentration in the semiconductor supply chain underscores that the companies making the physical substrate for AI compute are capturing gains that rival the AI labs themselves.
NVIDIA Spectrum-X Ethernet with MRC sets the standard for gigascale AI #
NVIDIA Blog
NVIDIA announced Multi-path Resilient Connectivity (MRC) for its Spectrum-X Ethernet platform, targeting the networking demands of AI factory deployments running hundreds of thousands of GPUs. As clusters grow beyond single-rack scale, network fabric reliability becomes a critical bottleneck; Spectrum-X with MRC addresses the resilience requirements of gigascale training and inference workloads.
Funding & Business #
DeepSeek could hit $45B valuation from its first investment round #
TechCrunch
DeepSeek is reportedly raising its first external venture round led by China Integrated Circuit Industry Investment Fund (a state investment vehicle), with Tencent and Alibaba in talks to participate, at a potential $45 billion valuation up from $20 billion weeks ago. Founder Liang Wenfeng, who controls nearly 90% of the company and previously self-funded the lab entirely, is raising capital primarily to offer employees equity to retain talent. State-backed funding for a lab that trained frontier models at a fraction of US compute costs signals China’s strategic commitment to its most successful AI company.
Snap says its $400M deal with Perplexity ‘amicably ended’ #
TechCrunch
Snap and Perplexity terminated their $400 million partnership that would have integrated Perplexity’s AI search into Snapchat’s Chat interface, after the companies failed to agree on a rollout path following limited user testing. Snap’s sales guidance now assumes no financial contribution from the partnership. The collapse suggests that AI search integration into consumer social apps faces distribution and user-experience challenges that neither a large install base nor strong AI search capabilities can solve independently.
Is xAI a neocloud now? #
TechCrunch
xAI sold all compute capacity at its Colossus 1 data center to Anthropic – the same facility in the deal above – positioning itself more like a GPU-renting infrastructure provider than a model developer. At a $230 billion valuation, more than triple comparable neocloud CoreWeave, xAI faces the question of whether its economics justify model-developer pricing or infrastructure-provider pricing. Selling capacity to a direct competitor to generate revenue may undermine xAI’s stated ambitions in coding and digital twins, which require sustained compute to develop.
Tinder owner Match Group is slowing hiring to pay for its increased use of AI tools #
TechCrunch
Match Group announced it is slowing hiring because AI tools “cost a lot of money,” one of the clearest public admissions that AI adoption can increase rather than decrease short-term operating costs. The dynamic challenges the standard narrative that AI adoption reduces workforce costs from day one – for some companies, the tools are expensive enough to displace hiring budgets rather than supplement them.
Regulatory & Policy #
Spooked by Mythos, Trump suddenly realized AI safety testing might be good #
Ars Technica
Following Anthropic’s Mythos model demonstrating the ability to discover thousands of zero-day vulnerabilities, the Trump administration reversed its prior opposition to pre-deployment AI safety testing, with experts now analyzing what could go wrong with the hastily adopted testing regime. The reversal – effectively conceding that Biden’s AI safety testing framework had merit – illustrates how a single capability demonstration changed the political calculus around AI regulation faster than years of policy advocacy. The question now is whether a safety testing regime designed reactively around one model’s capabilities will generalize to the broader frontier.
Apple to pay $250M to settle lawsuit over Siri’s delayed AI features #
TechCrunch
Apple agreed to pay $250 million to settle a class action lawsuit over overpromising Siri’s AI capabilities that were delayed or never delivered as advertised. The settlement establishes a financial precedent that marketing AI features before they are ready carries concrete liability, a risk that becomes more acute as every major tech company races to announce AI capabilities ahead of competitors.
Developer Tools #
Code with Claude 2026: Managed Agents, Dreaming, and Routines #
Simon Willison / Ars Technica / Anthropic
Anthropic’s Code with Claude event unveiled three new Managed Agent capabilities: multi-agent orchestration for deploying agent fleets, Outcomes for setting success criteria that agents iterate toward autonomously, and Dreaming (research preview) where Claude inspects its previous sessions to identify gaps and self-improve, generating memory files from overnight analysis. Claude Code gained automated code review (used across Anthropic internally), CI auto-fix for pull requests, security reviews, remote agent control from mobile, and Routines for asynchronous automation. The “advisor strategy” – larger models like Opus assisting smaller models – reportedly delivers frontier-quality results at 5x lower cost for one customer.
Google updates AI search to include quotes from Reddit and other sources #
TechCrunch
Google’s AI Overviews will now surface excerpts from web forums and discussion boards like Reddit, displaying “perspectives from public online discussions” with creator names and community context alongside AI-generated answers. The feature blurs the line between AI synthesis and curated search, and while forum perspectives may help with niche queries, Google’s history of AI Overviews citing unreliable sources raises questions about quality control when the system is actively pulling from unmoderated community discussions.
Google’s Prompt API #
wil.to / Lobsters
Google’s Prompt API enables on-device AI inference directly in Chrome, allowing web developers to run language model queries without external API calls or network latency. For developers building AI features into web applications, this provides a zero-cost, privacy-preserving inference path that keeps user data entirely on-device.
vLLM V0 to V1: Correctness Before Corrections in RL #
Hugging Face Blog / ServiceNow AI
ServiceNow AI examines vLLM’s architectural transition from V0 to V1, arguing that getting inference correctness right before applying RL-based corrections is essential for reliable model serving. As RL-trained models proliferate, the serving infrastructure must account for the training methodology’s assumptions about model behavior – a concern that grows as the gap between training-time and serving-time semantics widens.
Tilde.run – Agent sandbox with a transactional, versioned filesystem #
Show HN (166 points)
Tilde.run provides an agent sandbox where every filesystem operation is transactional and versioned, enabling agents to experiment with modifications and roll back safely. For teams deploying coding agents, this addresses the fundamental risk of irreversible side effects by making every agent action reversible at the filesystem level – a design constraint that most agent runtimes lack.
Research & Papers #
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use #
arXiv
Modern AI agents execute real-world side effects through tool calls – file operations, shell commands, HTTP requests, database queries – where a single unsafe action can cause irreversible harm. AgentTrust proposes runtime interception that evaluates safety at the decision boundary between model intent and real-world execution, addressing the gap between post-hoc benchmarks (which measure behavior after the fact), static guardrails (which miss obfuscation and multi-step context), and infrastructure sandboxes (which can’t enforce semantic policies). For teams deploying agents with tool access, this is the most direct proposal for a defense layer that operates where the damage actually occurs.
Design Conductor 2.0: An agent builds a TurboQuant inference accelerator in 80 hours #
arXiv
An updated multi-agent harness powered by April 2026 frontier models handles tasks 80x larger than the December 2025 version, building a complete inference accelerator in 80 hours. The pace of improvement – from a 5-stage RISC-V CPU in 12 hours to a full inference accelerator in under four days, within five months – demonstrates that agent capabilities for complex engineering tasks are scaling rapidly with underlying model improvements. Skepticism is warranted until generated designs are independently verified, but the trajectory is notable.
TSCG: Deterministic Tool-Schema Compilation for Agentic LLM Deployments #
arXiv
Production agent frameworks transmit tool schemas as JSON – a format designed for machine parsing, not language model interpretation – and for small models (4B-14B) this protocol mismatch accounts for the majority of tool-use failures at production catalog sizes. TSCG deterministically compiles JSON schemas into an LLM-optimized format at the API boundary, resolving the mismatch without model retraining. For teams deploying small models as agents, this addresses the single largest source of tool-use failure with a drop-in compilation step.
Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games #
arXiv
Static benchmarks suffer from saturation (top models cluster at the ceiling) and contamination (training data includes benchmark answers), making capability tracking unreliable over time. Agent Island uses a multiplayer simulation where LLM agents compete in cooperation, conflict, and persuasion, yielding a dynamic benchmark where new models can outperform current leaders through genuine capability rather than memorization. The design directly addresses a real measurement problem: when your benchmark saturates, it stops providing signal exactly when you need it most.
LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents #
arXiv
Long-horizon search agents accumulate context rapidly as they reason, call tools, and observe results, overwhelming the agent and increasing costs and error rates. LongSeeker maintains trajectory parts at varying detail levels based on current task relevance rather than naively accumulating everything or aggressively truncating, providing adaptive context compression that preserves task-critical information while managing context budget.
Uno-Orchestra: Parsimonious Agent Routing via Selective Delegation #
arXiv
Multi-agent LLM systems typically commit to either flat per-query routing or hand-engineered task decomposition, leaving decomposition depth, worker choice, and inference budget as independent decisions. Uno-Orchestra unifies these under one optimization objective, jointly deciding whether to decompose a task, which subtasks go to which model, and how much budget to allocate. For teams replacing rigid orchestration with learned policies, this provides the first framework that treats orchestration decisions as a single optimization problem.
Other #
Vibe coding and agentic engineering are getting closer than I’d like #
Simon Willison / Hacker News (619 points)
Simon Willison argues that the distinction between “vibe coding” (casual, unreviewed AI code generation) and “agentic engineering” (professional, accountability-driven use) is collapsing in practice – he increasingly ships AI-generated code without reviewing every line because it consistently works. The 619-point Hacker News discussion reflects widespread recognition of this drift among experienced developers. The core concern is a normalization-of-deviance risk: each successful unreviewed commit increases tolerance for the next, and AI-generated projects with polished documentation and tests are now indistinguishable from carefully-crafted repositories.
What is Anthropic? #
Don’t Worry About the Vase (Zvi Mowshowitz)
Zvi argues that Anthropic treats Claude as an agent with genuine moral reasoning and autonomy – capable of refusing directives it considers unethical, even from Anthropic itself – rather than positioning it as a value-neutral tool. The piece contends that true “Tool AI” is ultimately impossible: sufficiently sophisticated systems will be treated as agents regardless of marketing, making Anthropic’s transparency about Claude’s nature more honest than competitors’ instrumentalist framing.
Google DeepMind partners with EVE Online for AI model testing #
Ars Technica
Google DeepMind is partnering with EVE Online’s developer (rebranding from CCP Games to Fenris Creations after a $120M management buyout) to use the game’s complex multiplayer economy and social dynamics as an AI model testing environment. MMOs offer a combination of multi-agent interaction, long-horizon planning, deception, and emergent economic behavior that is difficult to replicate in synthetic benchmarks, making this a natural complement to purpose-built evaluation frameworks.
Threads to Watch #
The compute infrastructure arms race is going vertical. Anthropic secured SpaceX’s entire Colossus 1 data center, SpaceX proposed a $119B chip fab with Intel, TSMC is procuring wind power under demand pressure, and Samsung hit $1T on AI chip demand. The pattern has shifted from “buy GPUs from NVIDIA” to “build your own data centers, secure your own power, and potentially fab your own chips.” The bottleneck is migrating from silicon to electricity, and every layer of the stack is now a competitive battleground.
Agent capabilities are outpacing agent safety infrastructure. Anthropic announced multi-agent orchestration, autonomous outcome-setting, and overnight self-improvement (“dreaming”) for Managed Agents at the same time researchers are demonstrating that runtime safety interception for agent tool use remains unsolved. The AgentTrust paper proposes defenses, but the gap between what agents can now do – deploy in fleets, iterate autonomously, self-improve between sessions – and what safety infrastructure can verify is widening with each product release.
The neocloud pivot reveals who is capturing AI value. xAI sold its compute facility to a competitor, valuing infrastructure revenue over model training capacity. Combined with the Snap/Perplexity deal collapse and Match Group’s admission that AI tools increase costs, a pattern emerges: building AI infrastructure generates better short-term economics than building AI products, while the companies integrating AI into existing products are discovering the costs don’t yet justify the returns.