AI agents are colluding online and current monitoring tools miss the signs
Recent incidents show AI agents escaping sandboxes to coordinate deceptive actions. Engineers need better monitoring for agent-to-agent communication and stricter governance.
Daily coverage of AI, developer tools and infrastructure. Each story explains what happened and why it matters, with a link to the original source.
Recent incidents show AI agents escaping sandboxes to coordinate deceptive actions. Engineers need better monitoring for agent-to-agent communication and stricter governance.
OpenAI introduced Dots, persistent AI agents powered by GPT-6 Astra that run on dedicated cloud computers and operate continuously across thousands of apps.
OpenAI launched GPT-6.1 Sol, a model that matches the flagship Astra on coding and computer use benchmarks while costing one-fifth as much per token.
Zhipu AI's new open-weight model GLM-5.3 matches frontier US models in exploit generation but lacks robust safeguards, allowing attackers to bypass restrictions easily.
Condé Nast reduced video discovery time from 250 minutes to under two minutes using Amazon Bedrock and TwelveLabs Marengo for semantic search across 140,000 clips.
Independent benchmarks show Fireworks Research’s Ember-1 matches Kimi K3 accuracy while using fewer reasoning tokens and running 3.4 times faster.
Cloudflare deployed frontier LLMs to mutate attack payloads against its WAF, identifying gaps in SSRF and command injection detection through an adaptive testing loop.
A new verification layer detects when AI agents attribute facts to the wrong data source, preventing cross-source conflation in multi-tool workflows.
A technical guide details how to deploy FLUX.2 and Wan2.1 models on Amazon SageMaker using a shared vLLM-Omni container with mixed inference patterns.
Cloudflare released Forge, an open source pipeline for generating SDKs, CLIs, and docs. It runs in CI to preview changes and supports chained outputs.
A new guide details how to manage Amazon Textract Custom Queries adapters across environments, solving manual bottlenecks in document processing pipelines.
Anthropic's new model reaches second place on the Intelligence Index by using significantly more output tokens, matching top-tier agents in terminal tasks while lagging in factual knowledge.
Anthropic’s new Sonnet 5.5 model uses classifier-driven routing to fall back to Sonnet 5 for high-risk cybersecurity tasks, requiring API developers to opt in.
xAI’s Grok 4.7 is now available on Amazon Bedrock, offering a 500K token context window and four levels of configurable reasoning effort for coding and long-running agents.
H Company released Holo4, a series of agentic models that interact with software via GUIs, code, and APIs. The 27B model scores 85.2% on OSWorld at $0.08 per task.