Anthropic accuses Alibaba of the largest known distillation attack on Claude
Anthropic accused Alibaba of conducting the largest known distillation attack against Claude, using nearly 25,000 fraudulent accounts to extract capabilities over six weeks in a letter addressed to U.S. Senate leaders. OpenAI and Broadcom unveiled Jalapeno, OpenAI’s first custom inference chip, while Google added native computer use capabilities to Gemini 3.5 Flash. On the infrastructure side, Micron’s quarterly revenue quadrupled to $41.45 billion on AI memory demand, IBM demonstrated sub-1-nanometer chip technology, and Qualcomm agreed to acquire Modular to unify AI software across heterogeneous hardware.
Security #
Anthropic accuses Alibaba of illicitly extracting Claude AI model capabilities #
Anthropic / CNBC / Reuters / CybersecurityNews / Hacker News (443 points)
In a letter dated June 10, 2026 addressed to U.S. Senate Banking Committee Chair Tim Scott and Ranking Member Elizabeth Warren, Anthropic alleged that operators affiliated with Alibaba and its Qwen AI research division conducted a coordinated “adversarial distillation” campaign against Claude from April 22 to June 5, generating more than 28.8 million exchanges through nearly 25,000 fraudulent accounts. The operation specifically targeted Claude’s most commercially valuable capabilities, including software engineering and agentic reasoning from the Mythos Preview model. Anthropic warned that AI systems built through adversarial distillation often lack safety guardrails, framing the attack as both an intellectual property and a safety concern. The scale – the largest known distillation attack on any frontier model – sets a precedent for how model providers and policymakers think about cross-border AI capability extraction.
SoK: AI Secure Code Generation – Progress, Pitfalls, and Paths Forward #
arxiv:cs.AI
A systematization-of-knowledge survey covering prompting, fine-tuning, reinforcement learning, and agentic workflows for secure code generation finds that the field still lacks a systematic understanding of how these techniques improve security – or whether improvements hold across threat models and languages. For teams deploying code generation agents in production, the practical takeaway is that no single technique reliably produces secure code, and current evaluation frameworks do not adequately capture the security properties that matter in deployed systems.
AI Snitches Get Glitches: Towards Evading Agentic Surveillance #
arxiv:cs.AI / Hugging Face Daily Papers
This paper introduces and formalizes the risk of AI agents being repurposed as surveillance tools by employers or nation-states, then demonstrates techniques for evading such surveillance. The dual-use tension is direct: the same capabilities that make agents useful assistants (access to communications, data, APIs) make them effective surveillance infrastructure, and users may lack permission to control the surveilling agent’s actions. For anyone deploying enterprise agents, the implication is that agent access patterns need explicit privacy boundaries – not just for regulatory compliance, but because users will actively route around surveillance-capable agents that lack them.
How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring #
Hugging Face Daily Papers
Using 596 human-labeled completions, this paper compares safety classifier judges against prompted chat-model judges and finds they fail in opposite directions: classifiers are overconfident on boundary cases while chat-model judges are systematically lenient. The two judge families are also differentially vulnerable to adversarial attacks. For teams relying on automated safety evaluation – which is nearly everyone reporting attack-success rates – the finding means that published ASR numbers are substantially noisier than assumed, and the choice of judge architecture changes which attacks appear to succeed.
Infrastructure #
OpenAI and Broadcom unveil Jalapeno, an LLM-optimized inference chip #
OpenAI / TechCrunch / Ars Technica / Hacker News (577 points)
OpenAI and Broadcom announced Jalapeno, OpenAI’s first custom silicon designed specifically for LLM inference workloads. The chip targets the inference-specific bottlenecks – memory bandwidth, attention computation, and batch scheduling – that general-purpose GPUs handle suboptimally, following the path Google blazed with TPUs and Amazon with Trainium. For the broader AI infrastructure market, Jalapeno signals that the largest model providers now view custom silicon as table stakes for inference economics: as inference volume scales faster than training compute, the per-token cost structure becomes the competitive battleground.
IBM demonstrates sub-1-nanometer chip technology #
MIT Technology Review / Ars Technica
IBM built a prototype chip with approximately 100 billion transistors on a fingernail-sized area using its new “nanostack” transistor architecture, doubling the density of its previous state-of-the-art 2nm technology from 2021. The design could extend Moore’s Law by another decade, delivering either faster performance or improved energy efficiency at each subsequent node. For AI infrastructure, denser transistors translate directly into more compute per watt – the metric that determines the economics of both training and inference at scale.
Micron quarterly revenue quadruples to $41.45 billion as AI memory demand surges #
TechCrunch
Micron Technology reported quarterly revenue of $41.45 billion (up from roughly $10 billion year-over-year) with profits surging to $28.2 billion, driven by the AI-induced memory chip shortage predicted to persist through 2027. Shares jumped 13% after earnings. The company also secured a strategic supply deal with Anthropic, signaling that memory capacity is now a bottleneck significant enough for AI labs to negotiate direct supplier relationships rather than buying through standard channels.
I/O Design Challenges Grow in AI Data Centers and HPC Clusters #
Semiconductor Engineering
Physical I/Os are becoming a chokepoint for high-performance AI chips as interconnect protocols push bandwidth limits, requiring design tradeoffs between performance, power, and reliability. As AI accelerators get faster, the pins and SerDes links connecting them to memory and networking become the binding constraint – a shift from the compute-bound bottlenecks that dominated the previous generation of AI hardware design.
Model Releases #
Introducing computer use in Gemini 3.5 Flash #
Google / Google DeepMind / Hacker News (222 points)
Google integrated computer use as a native built-in tool in Gemini 3.5 Flash, enabling agents to see, reason about, and take action across browser, mobile, and desktop environments. The capability was previously available only in a standalone Gemini 2.5 model but is now part of the main Flash offering, with targeted adversarial training to mitigate prompt injection and optional enterprise safeguards requiring user confirmation for sensitive actions. Google pitches the feature for long-horizon automation tasks like continuous software testing and enterprise knowledge work. The integration into Flash rather than a premium tier makes computer use a commodity capability for agent builders, shifting the competitive question from “can your model use a computer” to “how reliably and cheaply.”
Funding & Business #
Qualcomm to acquire Modular #
Modular / Reuters / Hacker News (577 points)
Qualcomm agreed to acquire Modular, the AI software infrastructure company founded by Chris Lattner (creator of LLVM and Swift), in a deal expected to close in the second half of 2026. Modular’s platform enables AI model deployment across heterogeneous hardware – CPUs, GPUs, NPUs, and custom ASICs – without code rewrites, addressing the fragmentation problem that forces developers to optimize separately for each accelerator. The acquisition gives Qualcomm a “write once, run anywhere” software layer for AI as the company pushes into data center and edge AI markets alongside its established mobile position. For teams deploying models across diverse hardware, the consolidation of Modular’s cross-platform compiler technology into a chip vendor raises questions about whether the platform remains truly hardware-neutral.
Cerebras stock plunges after first public earnings report #
TechCrunch
Cerebras Systems’ stock fell nearly 20% following its first earnings report as a public company, despite beating expectations with $193 million in quarterly revenue and a narrowed net loss. The sell-off was triggered by full-year gross margin guidance of 38-41%, well below the 47% margin in Q1, which CEO Andrew Feldman attributed to a strategic decision to temporarily rent systems back from a customer while building data center capacity. The market reaction illustrates the gap between AI hardware companies’ revenue growth narratives and the margin compression that comes with building out physical infrastructure – a dynamic that affects every AI chipmaker transitioning from chip sales to compute-as-a-service.
Agility Robotics plans to go public via SPAC in a $2.5B deal #
TechCrunch
Agility Robotics is merging with Churchill Capital Corp XI in a deal valuing the humanoid robotics company at approximately $2.5 billion, generating over $620 million in proceeds including $200 million from institutional investors. The company plans to scale production of its Digit v5 robot and fulfill $300+ million in existing orders from Toyota, GXO, and Mercado Libre. The SPAC route and customer roster suggest the humanoid robotics market is reaching the point where unit economics can be demonstrated to public market investors, though whether the valuation holds depends on production scaling rather than capability demos.
AI researchers continue to leave Google for its rivals #
TechCrunch
Top AI researchers Jonas Adler and Alexander Pritzel are departing Google for Anthropic, following recent high-profile exits by Nobel laureate John Jumper (also to Anthropic) and Noam Shazeer (to OpenAI). The departures from Google’s Gemini model team reflect how approaching IPOs at Anthropic and OpenAI create equity-based recruitment leverage that established companies cannot match. The talent pipeline from Google to its competitors has become a structural feature of the AI industry rather than an anomaly.
Companies are scrambling to stop employees from maxing out AI budgets #
TechCrunch
After initially encouraging employees to maximize AI usage through budgets and leaderboards, companies are now reversing course as costs spiral unpredictably. Consulting firms like Accenture are implementing restrictions to prevent employees from using expensive AI models for routine tasks like converting PDFs to slides. The shift from “tokenmaxxing” to “token rationing” reveals a fundamental challenge in enterprise AI adoption: organizations cannot yet distinguish high-value AI usage from wasteful consumption, and the cost-per-query model makes every interaction a budget decision.
Research & Papers #
Constraint Tax in Open-Weight LLMs: Tool Calling Suppression Under Structured Output Constraints #
Hugging Face Daily Papers
When tool calling and JSON Schema constraints are simultaneously enabled, multiple open-weight models cease invoking tools despite maintaining high schema compliance – a phenomenon the authors call “tool suppression.” The finding is directly relevant to anyone building agent systems on open-weight models: the two capabilities most critical for agents (structured output and tool use) interfere with each other in ways that benchmark scores on either capability alone would not predict. The practical implication is that agent builders using open-weight models need to test the joint deployment configuration, not each capability in isolation.
Quantization Inflates Reasoning: Token Inflation as a Hidden Cost of Low-Bit Reasoning Models #
arxiv:cs.AI
Low-bit post-training quantization introduces a hidden inference cost: quantized reasoning models generate longer chains of thought even when they still answer correctly, meaning the per-token savings from quantization are partially offset by generating more tokens. The effect spans mathematical reasoning, code generation, and scientific questions. For teams deploying quantized models to reduce inference costs, the total compute savings may be substantially smaller than the per-token cost reduction suggests – a factor that should change how quantization’s ROI is calculated.
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs #
Google Research
Google researchers found that enabling reasoning traces significantly improves LLMs’ ability to recall simple factual information, even when complex problem-solving is not necessary. The mechanism works through a “computational buffer” that refines internal states and a “factual priming” effect where generating related facts during reasoning primes correct answer retrieval. However, the mechanism is fragile: hallucinated intermediate facts substantially degrade final accuracy. The implication for production systems is that reasoning-enabled models are better fact retrieval engines, but only when the intermediate reasoning is factually grounded.
Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning #
arxiv:cs.AI
This paper identifies “cliff tokens” – single tokens where the potential for a correct answer drops sharply, triggering cascading reasoning failures. Unlike prior work that analyzes failure at the step or sentence level, cliff tokens pinpoint the exact token where the model commits to a wrong path. For teams building reliable reasoning systems, this provides a diagnostic primitive: monitoring token-level potential during generation could enable early intervention before a reasoning trace goes irrecoverably wrong.
TRUSTMEM: Learning Trustworthy Memory Consolidation for LLM Agents #
arxiv:cs.AI
Current memory-augmented LLM agents actively update external memory through write, revise, and delete operations, but these updates may omit important information, corrupt existing memory, or introduce hallucinated content that becomes a persistent system-state failure. TRUSTMEM proposes a learned consolidation mechanism that makes memory updates more reliable. For agent builders, the contribution addresses a failure mode that gets worse over time: small memory errors compound across interactions, and once stored, incorrect information becomes load-bearing context for all future decisions.
Regulatory & Policy #
Europe is pushing back on Washington’s chip war #
TechCrunch
Dutch Trade Minister Sjoerd Sjoerdsma visited Washington to lobby Congress against the MATCH Act, which would extend U.S. export controls to ASML’s deep ultraviolet lithography machines – decade-old technology that China can currently purchase. ASML, Europe’s most valuable company and the sole global manufacturer of these machines, would be significantly impacted as China accounts for 19% of its sales. The dispute crystallizes the tension between U.S. national security strategy and European economic interests, with chip export controls becoming the most concrete point of transatlantic AI policy disagreement.
Developer Tools #
How To Give Your Agent Memory #
LangChain Blog
LangChain published a practical guide to implementing agent memory, covering techniques for enabling agents to retain and utilize information across conversations. For teams building on LangChain’s agent framework, the post provides implementation patterns for the memory layer that separates stateless chat completions from agents that accumulate context over time – a capability gap that most agent frameworks leave to the builder to solve.
Figma adds code layers, animations, and AI features in new update #
TechCrunch
Figma’s latest update introduces AI-powered capabilities including text-prompt creation of custom plug-ins, an AI assistant with external tool connections to Notion and GitHub, and AI-generated shader effects. The platform now supports AI agents that execute repeatable tasks within the collaborative canvas. The AI features position Figma as another design tool where the workflow is shifting from manual creation to prompt-driven generation with human curation.
Open Source #
RubyLLM: A Ruby framework for all major AI providers #
Hacker News (392 points)
RubyLLM is a unified Ruby framework for interacting with all major AI providers, receiving significant community attention on Hacker News. The framework fills a gap in the Ruby ecosystem where Python and TypeScript have dominated AI tooling, providing Ruby developers with first-class abstractions for multi-provider LLM integration. The strong community response suggests pent-up demand for AI development tools outside the Python-centric ecosystem that currently dominates.
Threads to Watch #
Model distillation is becoming an adversarial battleground. Anthropic’s accusation that Alibaba ran a coordinated 28.8-million-exchange distillation campaign – targeting specific capabilities like agentic reasoning and software engineering – elevates model extraction from a theoretical concern to a documented, large-scale attack vector. Combined with the SoK survey showing that secure code generation techniques remain unreliable, the emerging picture is that protecting model capabilities requires defenses at the API access layer (rate limiting, behavioral detection) rather than at the model layer, because the model cannot distinguish legitimate heavy use from systematic capability extraction.
Custom inference silicon is now a multi-player race. OpenAI’s Jalapeno joins Google’s TPUs, Amazon’s Trainium, and Microsoft’s Maia in a market where every hyperscaler and major model provider is building inference-specific hardware. IBM’s sub-1nm demonstration and Micron’s $41B quarter show the supporting semiconductor ecosystem scaling to meet demand. The strategic implication: inference cost, not training cost, is becoming the primary economic driver as deployed AI workloads grow faster than model training cycles.
Open-weight agent deployment has hidden capability interactions. The Constraint Tax paper demonstrates that tool calling and structured output – the two capabilities most critical for agents – suppress each other in open-weight models, while quantization inflates reasoning token counts, partially negating cost savings. These findings suggest that benchmarking individual capabilities in isolation systematically overestimates real-world agent performance, and that deployment-configuration testing needs to become a standard practice.
Sources Unavailable Today #
These sources could not be fetched today. Links point to their homepages so you can check them directly.
- Weights & Biases: Fully Connected — scrape: content not extractable