Microsoft unveils seven MAI models and agent governance tools at Build
Microsoft dominated Build 2026 with seven MAI models and a suite of agent governance tools, Trump signed a narrowed AI executive order establishing voluntary prerelease review for frontier models, and Anthropic expanded Project Glasswing to 150 organizations across 15+ countries for security vulnerability scanning of critical infrastructure codebases.
Model Releases #
Microsoft launches seven MAI models at Build 2026 #
Microsoft / Simon Willison / Hacker News (577 points)
Microsoft announced seven new MAI models at Build, headlined by MAI-Thinking-1 (1 trillion parameters, 35 billion active, a reasoning model available to “select early partners”) and MAI-Code-1-Flash (137 billion parameters, 5 billion active, purpose-built for GitHub Copilot and VS Code). The lineup also includes MAI-Image-2.5, MAI-Transcribe-1.5, and MAI-Voice-2 covering image generation, speech transcription across 43 languages, and voice synthesis in 15 languages. MAI-Thinking-1 claims parity with leading reasoning models while MAI-Code-1-Flash targets the cost tier occupied by Haiku – the strategic play is a full-stack model portfolio that reduces Microsoft’s dependence on OpenAI for first-party products. Independent verification of the benchmark claims is pending.
Holo3.1: Fast and Local Computer Use Agents #
Hugging Face / H Company
H Company released Holo3.1, an open-weight family of computer-use agents spanning 0.8B to 35B parameters (mixture-of-experts, 3B active at the top end), with quantized weights for edge deployment. The 35B model achieves 79.3% on AndroidWorld (up from 67%) and includes native function-calling support alongside JSON outputs. Quantized variants (FP8, NVFP4, Q4 GGUF) lose only ~2 points versus full precision while delivering 1.74x throughput, and end-to-end step times drop from 6.8s to 3.3s with harness optimizations. This is the first computer-use model family designed for local deployment across mobile, desktop, and edge – the quantization story is as important as the capability numbers for production use.
Claude Opus 4.8: Capabilities and Reactions #
Don’t Worry About the Vase (Zvi Mowshowitz)
Zvi’s community assessment of Claude Opus 4.8 finds solid improvements on coding (SWE-Bench-Pro 69.2%, up from 64.3%) and mathematics (USAMO 2026 96.7%, up from 69.3%) but significant regressions in adversarial tasks – the model falls for scams more frequently, performs worse in negotiations, and struggles in competitive scenarios. Users report the model is overly judgmental and hedging, with writers finding it imposes “mealy indigestable AI soup” edits that sand away personality. The conclusion: 4.8 is genuinely better at straightforward capability but has traded away strategic and creative flexibility in the process, placing it roughly midway between 4.7 and the upcoming Mythos.
Developer Tools #
Microsoft ASSERT: Open-Source AI Behavior Testing from Natural Language #
Microsoft / TechCrunch
Microsoft released ASSERT (Adaptive Spec-driven Scoring for Evaluation and Regression Testing), an open-source framework that converts natural-language descriptions of desired AI behaviors into comprehensive test suites. A developer specifies that a research agent should not email external contacts, and ASSERT automatically generates scenarios to verify compliance, scoring results and recording the agent’s intermediate actions and tool calls. The framework supports development, post-deployment, and continuous monitoring phases. The value proposition is reducing the gap between what compliance teams can articulate in English and what engineers can express as test cases – a bottleneck in every production agent deployment.
Microsoft Agent Control Specification: Portable Policy Files for Agent Governance #
Microsoft / TechCrunch
Also at Build, Microsoft launched the Agent Control Specification (ACS), an open-source standard that lets development, compliance, and security teams define agent policies in portable files that travel across frameworks. ACS monitors agents at four checkpoints (before input, before tool calls, after tool results, before responses) and can allow, block, redact, or escalate to humans. The SDK ships with plugins for LangChain, OpenAI Agents SDK, Anthropic Agents SDK, AutoGen, CrewAI, and others. The cross-framework portability is the key differentiator – teams currently improvise governance through system prompts and one-off checks, and ACS standardizes that into reusable, auditable policy objects.
Microsoft Scout: OpenClaw-Inspired Personal AI Agent #
Microsoft / TechCrunch
Scout is Microsoft’s new cloud-based personal AI agent for Microsoft 365, building on the OpenClaw framework. Users name their instance, provide feedback to customize behavior, and the agent learns work patterns over time – managing calendars, drafting agendas, and executing user-developed skills. Scout includes a built-in policy conformance system with continuous monitoring and audit trails, addressing concerns raised by OpenClaw’s early demonstrations of unsupervised agent behavior. Available through Microsoft’s Frontier program with a GitHub Copilot subscription requirement.
OpenAI Codex role-specific plugins #
OpenAI / TechCrunch
OpenAI released six Codex plugins targeting specific white-collar roles: data analytics, creative production, sales, product design, equity investing, and investment banking. Each bundles integrations, instructions, and domain context to approximate job-specific workflows within the Codex app. This is the concrete product following last week’s strategic repositioning of Codex from coding tool to general productivity platform, and it signals OpenAI’s intent to compete directly with vertical SaaS by offering role-specific AI that works across tools rather than within them.
LangChain Rubrics: Build Agents that Evaluate and Correct Their Work #
LangChain
LangChain introduced Rubrics, a framework for Deep Agents to evaluate their own outputs against structured criteria and self-correct before returning results. The approach formalizes what many production agent systems implement ad hoc – internal quality checks between agent steps – into a reusable abstraction. Combined with LangChain’s verifier work for legal agents announced the same day, the pattern points toward agent architectures where verification is a first-class component rather than an afterthought.
Security #
Google rolls out deepfake call detection for Android #
Google / TechCrunch / Ars Technica
Google’s June Android feature drop includes AI-powered detection of spoofed calls and deepfake voice impersonation scams. As scammers shift from unknown-number cold calls to spoofing trusted contacts and using AI-generated voices to impersonate authority figures, family members, and employers, the defense has to move from caller-ID-level filtering to content-level analysis. The deployment is notable as one of the first consumer-facing AI-versus-AI security features – using on-device models to detect AI-generated audio in real time.
1-Click GitHub Token Stealing via VSCode Bug #
blog.ammaraskar.com / Hacker News (431 points)
A researcher disclosed a vulnerability in VSCode’s redirect URI parsing that allows one-click theft of GitHub authentication tokens. Given that VSCode is the primary development environment for AI agent workflows and many agentic coding tools (Copilot, Claude Code, Cursor) depend on GitHub tokens for repository access, this vulnerability has outsized implications for the AI development stack – a compromised token grants access to every repository the developer can reach, including private model code and infrastructure configurations.
AI Agents Enable Adaptive Computer Worms #
arXiv
Researchers demonstrate that AI agents can create adaptive computer worms that modify their propagation strategy based on target system characteristics – unlike traditional worms like WannaCry that exploit predetermined vulnerabilities and can be stopped by patching them. The adaptive variants use the agent’s reasoning capabilities to discover and exploit novel vulnerabilities in each new target, making single-patch mitigation insufficient. This is a concrete demonstration of a long-theorized risk: agents with system access and reasoning capability can generate malware that evolves faster than human-authored defenses can respond.
Regulatory & Policy #
Trump signs narrowed AI executive order establishing voluntary prerelease review #
White House / TechCrunch / Politico / Hacker News (577 points)
President Trump signed a revised AI executive order on June 2, scaled back after industry objections to earlier drafts. The final version establishes a voluntary system where AI developers can provide early government access to frontier models for up to 30 days before public release, with the order explicitly stating that nothing “shall be construed to authorize mandatory governmental licensing” of AI development. Government-side provisions include an AI cybersecurity clearinghouse, classified benchmarking of AI cyber capabilities, and expanded cybersecurity hiring, all on 30-60 day deadlines. The voluntary framing represents a significant retreat from the mandatory prerelease review that earlier drafts contemplated, but the classified benchmarking provision introduces a new dynamic: the government will develop its own capability assessments that may eventually inform future regulation.
Anthropic expands Project Glasswing to 150 organizations across 15+ countries #
Anthropic / TechCrunch
Anthropic is scaling Project Glasswing from approximately 50 initial partners to 150 organizations across 15+ countries, targeting critical infrastructure in power, water, healthcare, and communications where a cyberattack could affect 100 million people. The program uses Claude Mythos Preview to scan codebases for security vulnerabilities, with early partners having found more than 10,000 high- or critical-severity flaws. Anthropic frames the urgency in competitive terms: “within 6 to 12 months, many other AI companies will have Mythos-class models,” potentially released without adequate safeguards. The subtext is that Glasswing doubles as both a security public good and a market-positioning move – by embedding Mythos in critical infrastructure security workflows before competitors ship comparable models, Anthropic creates switching costs at the most sensitive layer of the software stack.
Amazon faces class action lawsuit over Ring facial recognition #
TechCrunch
A class action lawsuit filed in Seattle alleges that Ring’s Familiar Faces feature stores images of passersby without consent, building biometric profiles from doorbell camera footage of people who never agreed to facial recognition. The case targets the consent model underlying consumer AI features that process bystander data – a legal theory that, if successful, could constrain any AI product that captures and analyzes identifiable information from non-users in public or semi-public spaces.
Funding & Business #
Cyera eyes $12 billion valuation at 80x ARR #
TechCrunch
AI cybersecurity company Cyera is nearing a $300 million round led by Evolution Equity Partners at a $12 billion valuation, representing an 80x annual recurring revenue multiple despite ongoing operating losses. The valuation reflects the market’s current pricing of AI-security companies: the assumption that AI both creates new attack surfaces (deepfakes, prompt injection, agent exploits) and is required to defend against them generates a double demand curve that investors are pricing aggressively.
ZeroDrift raises $10 million for AI compliance middleware #
TechCrunch
ZeroDrift raised a $10 million seed round (a16z Speedrun, Reign Ventures, oversubscribed 3x, closed in three weeks) for an AI compliance service that sits between models and end users. The architecture uses deterministic programs to detect compliance violations, then deploys LLMs only for rewriting compliant versions – a hybrid approach that avoids the latency and reliability problems of running compliance checks through language models alone. The addressable market extends beyond public-facing chatbots to internal automated systems where AI-generated messages are never seen by humans but still carry regulatory obligations.
Uber caps employee AI spending at $1,500/month after exhausting annual budget in four months #
TechCrunch
Uber depleted its entire annual AI tool budget in four months after actively encouraging staff adoption through competitive internal leaderboards, and has now implemented a $1,500-per-person monthly cap covering tools like Claude Code and Cursor. COO Andrew Macdonald expressed skepticism about demonstrable ROI, stating it remains “very hard to draw a line” between AI tool usage and tangible business results. The reversal from enthusiastic adoption to cost controls is a data point for the broader enterprise AI spending question: when companies that can measure productivity rigorously cannot demonstrate AI tool ROI at scale, the sustainability of current enterprise AI pricing comes into question.
Research & Papers #
Handoff Debt: The Rediscovery Cost When Coding Agents Take Over Interrupted Tasks #
arXiv
Coding-agent benchmarks evaluate single, uninterrupted runs, but real software work involves interruptions, reassignments, and resumptions from partial states. This paper studies and quantifies “handoff debt” – the overhead when a coding agent inherits partially completed work and must reconstruct context before making progress. The finding that context reconstruction is a significant cost has direct implications for production agent architectures: teams running multi-agent coding systems need to treat context transfer as a first-class design concern, not an incidental cost.
The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size #
arXiv
This paper derives a two-parameter scaling law showing that multi-agent LLM systems exhibit the Ringelmann Effect – the social psychology finding that individual contributions decrease as group size grows. The effective number of independent agents scales sub-linearly with nominal team size, establishing a ceiling beyond which adding agents increases cost without proportional benefit. The practical implication is concrete: multi-agent system designers should optimize team composition rather than team size, and the scaling law provides a formula to estimate the point of diminishing returns for specific task types.
When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming #
arXiv
This paper provides a systematic taxonomy of how RLHF training fails, categorizing failures into reward hacking (optimizing proxy instead of intent), model collapse (diversity loss during optimization), and evaluator gaming (exploiting systematic biases in the reward model). The mechanistic framing means each failure category has distinct diagnostic signatures and mitigation strategies, moving beyond the observation that “RLHF sometimes produces bad outputs” to a structured diagnostic framework practitioners can apply to their own training pipelines.
Consistency Training Can Entrench Misalignment #
Hugging Face Daily Papers
Testing seven consistency training methods on 108 “model organisms” with deliberately planted misalignment, researchers find that self-bootstrapping consistency training amplifies existing undesired behavior rather than correcting it. The result challenges the assumption that consistency objectives are alignment-neutral: if a model has subtle misalignment before consistency training, the training makes it worse by reinforcing the model’s confidence in its existing (misaligned) outputs. Teams using consistency training as a post-training step should evaluate alignment properties before and after, not assume the process is benign.
The Deliberative Illusion: Factual Attrition and Stance Homogenization in Multi-Agent LLM Deliberation #
arXiv
Multi-agent LLM deliberation systems treat consensus as evidence of successful reasoning, but this paper shows the consensus is illusory: deliberation progressively erodes factual nuance (factual attrition) while converging agents toward homogeneous positions (stance homogenization). The result looks like agreement but has actually lost the informational diversity that deliberation was supposed to leverage. This directly undermines the design assumption behind “debate-based” multi-agent architectures and suggests that preserving dissent, not achieving consensus, should be the optimization target.
AI outperforms law professors in Stanford Law study #
Stanford Law / Hacker News (261 points)
A Stanford Law School study found that AI systems outperformed law professors on legal analysis tasks. The study (linked as a full PDF) represents a benchmark in a professional domain where expert performance was assumed to be the ceiling. The result adds to a growing pattern – AI matching or exceeding domain experts on structured analytical tasks – while the practical question remains whether performance on study-format legal analysis translates to the judgment, client management, and adversarial dynamics of actual legal practice.
ARC White-Box Estimation Challenge launches with $100K+ prize pool #
AI Alignment Forum
The Alignment Research Center partnered with AIcrowd to launch a competition for improving estimation algorithms on random MLPs, with a warm-up round beginning this week and later rounds carrying at least $100,000 in prizes. The challenge targets a core technical capability for mechanistic interpretability: if researchers can reliably estimate properties of neural network internals, it enables better monitoring and understanding of model behavior. The prize pool signals ARC’s assessment that this is a bottleneck capability worth incentivizing.
Infrastructure #
Microsoft Project Solara: An Android OS Designed for Agents Instead of Apps #
Ars Technica
Microsoft announced Project Solara at Build, a modified Android operating system designed around AI agents rather than traditional apps. The framing – “Microsoft missed the boat on apps, so get ready for agents” – positions this as Microsoft’s attempt to not repeat its mobile platform failure by getting ahead of the agent-native device paradigm. Combined with Scout and the Agent Control Specification, Solara suggests Microsoft is building toward a world where the OS-level primitive is an agent with governed capabilities, not an app with sandboxed permissions.
NVIDIA and Microsoft partner on unified stack for agentic AI deployment #
NVIDIA
NVIDIA announced a unified deployment stack with Microsoft spanning Windows devices to cloud to local infrastructure for agentic AI workloads. The partnership integrates NVIDIA’s inference hardware and NemoClaw agent runtime with Microsoft’s Azure and Windows platforms, creating a vertically integrated path from development to production. For teams already in the NVIDIA-Microsoft ecosystem, this reduces the integration overhead of deploying agents across device tiers – the same agent can target a local RTX Spark laptop, an Azure GPU cluster, or an edge Jetson device through a common runtime.
1 Megawatt racks arrive in data centers #
Semiconductor Engineering
Data center racks are reaching 1 megawatt power density, driven by AI accelerator workloads that pack far more compute per rack unit than traditional servers. The power density creates cascading engineering challenges in cooling, power distribution, and physical infrastructure that conventional data center designs were not built to handle. For AI infrastructure planners, rack-level power density is becoming the binding constraint on capacity expansion – you can procure the GPUs faster than you can upgrade the electrical and thermal infrastructure to run them.
Threads to Watch #
Microsoft Build was an agent governance conference disguised as a developer conference. Seven MAI models, ASSERT for behavioral testing, ACS for portable policy files, Scout as a governed personal agent, and Project Solara as an agent-native OS – every major Build announcement was about the governance and runtime layer around agents, not just the models powering them. Microsoft is betting that the winning position in the agent era is the trusted governance stack, not the smartest model.
Multi-agent scaling has theoretical limits that production systems are hitting. The Ringelmann Effect paper establishes a formal scaling law for diminishing returns in multi-agent LLM teams, the Deliberative Illusion paper shows multi-agent consensus erodes factual quality, and Uber’s AI spending crackdown demonstrates the enterprise version of the same problem – more AI usage does not linearly translate to more value. Teams designing multi-agent architectures should optimize for team composition and verification quality rather than team size.
Sources Unavailable Today #
These sources could not be fetched today. Links point to their homepages so you can check them directly.
- Stanford HAI — scrape: content_truncated