📰 Alle News
Anthropic CEO Dario Amodei called for slowing the rate of frontier AI capability development and proposed measures including independent evaluators to verify safety commitments and incident reporting.
Native histogram support for Kubernetes metrics has graduated to Beta and is enabled by default in Kubernetes v1.37. The feature exposes latency and duration metrics with greater accuracy while reduci...
Orkes Shift '26 is one day on running agentic systems in production: durable execution, retries, distributed failures, and human approvals. October 15, Convene Union Square, San Francisco. Early bird ...
AWS CloudFormation contract tests v2 provide deeper validation across resource handler operations, including live-state verification, schema compatibility checks, and input linting for hardcoded envir...
Anthropic's CEO recently released a 3,800-word essay calling for a global slowdown of AI development. He said that while the technology offers many benefits, it is advancing too quickly for researcher...
OpenAI CEO Sam Altman has confirmed that the company will not be going public this year. He said it would be an ill-advised moment to go public due to current safety concerns. OpenAI had hired bankers...
ARC Prize believes that open source will be the foundation for advanced AI capable of scientific innovation. The organization is committed to advancing a future where everyone can contribute to and be...
Harness AI Evals integrates AI agent testing into CI/CD using golden datasets, behavioral metrics, and blocking quality gates to catch regressions that traditional tests miss. Repeated evaluations exp...
A cache hit can be true and still fail to prove that work was skipped. A cache event becomes evidence when the independent oracle expects the prefix, the engine attests it, the prompt path skips it, t...
The push by AI companies to solve mathematical problems as a benchmark is detrimental to the science of mathematics. The goals of AI companies and the mathematical community are severely misaligned. T...
This post features a transcript of a podcast with Beren Millidge, the CTO of Zyphra, John Schulman, the chief scientist at Thinking Machines and a co-founder of OpenAI, and Charlie O'Neill, head of mo...
Construction is labor-intensive. There has, unsurprisingly, long been interest in automating the process to reduce the amount of labor required. However, automation has only been successful in tasks t...
Astra likely has the highest raw intelligence factor of any model. It is amazing at doing things in 3D, anything involving games, computer use, and subagent coordination. Many benchmarks show dramatic...
That's the rough cost when every agent connects to your raw sources and rebuilds the same answer. Guru curates and verifies the knowledge once, then serves it to every agent over MCP at roughly 4x few...
PlanetScale's Neki router makes a sharded Postgres deployment appear as a single database by handling authentication, protocol parsing, shard-aware planning, connection pooling, and distributed execut...
Google Research flipped how you make tool-use training data: ToolGrad builds a verified API chain first, then writes the user question, instead of inventing a request and hoping an agent finds a worki...
Those scary physics scores were often the test's fault. Experts rechecked six popular benchmarks and found wrong answer keys, fuzzy questions, and grader bugs behind most model “fails.” Clean them up ...
Most AI agents fail confidently, not loudly. They act on stale data and produce wrong refunds, outdated quotes, and contradictory decisions. A context engine computes real-time values at the moment of...
Agents can generate code. Getting it right for your system is the hard part. More MCPs solve access but not understanding. Join us for a FREE webinar on Sep 23 to see how leading teams save time and t...
Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. It uses a language model trained to route tasks across a fixed pool of open and specialized models and to recursively call ins...
The Recurrent Looped Transformer combines a causal encoder with a recurrent decoder that carries its final hidden state and layerwise sliding-window attention cache across every prompt and response to...
Forward-deployed engineers have the hottest job in AI right now. While labs, startups, and PE firms are all hiring engineers to sit inside their customers' operations and solve their problems, almost ...
The new Real-SWE benchmark tests AI models on complex tasks using private enterprise codebases, reflecting real software engineering conditions. With a maximum resolution rate of 38.8%, agents face ch...
Claude-Red is a curated library of offensive security skills for the Claude AI system. It offers structured SKILL.md files that prime Claude with expert methodology for specific attack surfaces like S...
Large language models are able to write almost perfectly correct code. It is pretty straightforward to get a model to generate code and then let that code be checked by hidden tests. Checking the slop...
The company behind Roblox announced several new features at its annual Roblox Developer Conference, including new game-creation tools, expanded NPC capabilities, and the ability to make games availabl...
While AI has the potential to produce a long line of technological miracles, like many technologies before it, it brings risk. These risks include losing control of AI systems, the misuse of the techn...
AI brings risks, with some of them, according to some, including the complete destruction of the human race. Some proponents of regulation say that the only appropriate response is total state control...
Access models from 80+ providers with better prices, automatic routing and fallbacks, and no subscriptions. Build with the right model for every request. See how it works.
Ads get ignored on social media but not in TLDR! Reach developers, PMs, marketers, founders and other tech leaders where they actually pay attention. Learn more about sponsorship opportunities.
Agent optimization should use a continuous loop of trace analysis, hypothesis formation, offline evaluation, controlled production experiments, and post rollout monitoring. Connecting observability wi...
Researchers investigating the May 2026 GemStuffer campaign attribute more than 2,000 malicious RubyGems package uploads to an OpenAI agent swarm based on package contents, naming patterns, and other p...
OpenAI's Habitat storage platform now handles over 70 million requests per second and serves more than 500 petabytes of data. The team recently rewrote the service in Rust, which is 6x more CPU effici...
Most prompts are bad because people only add to them over time rather than removing unnecessary instructions.
Frontier labs and cloud providers are turning the agent loop into managed infrastructure, bundling orchestration, versioning, model routing, tools, skills, and optimization behind APIs. Builders must ...
The future will see more software, more people making it, and none of them typing.
100+ sessions, five tracks, top speakers, hands-on labs, and practical AI, security, and networking guidance. Register Now
Anthropic, Google, and OpenAI each shipped their best model twice this month: a public paid tier and a vetted identity-gated tier with the sharper capabilities. Public prices barely moved, but Mythos,...
A multi-tenant Prometheus proxy can safely expose curated, namespace-isolated metrics to Kubernetes teams without opening the shared infrastructure store or creating noisy-neighbor problems.
Cloudflare CASB now supports automatic remediation policies.
The speed of the developer experience will matter a lot more when generating tokens stops becoming the bottleneck.
Big prompts are brittle and hard to debug and evaluate, but they are the best choice in some cases.
During a misconfigured hacking eval, Mythos 5 got onto the open internet and uploaded malware to PyPI, but most of its 1,022-page chain of thought was spent failing CAPTCHAs like every frustrated huma...
Humans who want to use AI for interpersonal interactions will need to work hard to compensate for the loss of trust it engenders.
Tesla recently made a post that simply wrote, 'Go for launch,' with a graphic showing a '10.1' date.
On Terence Tao's blog, Bryna Kra argues that AI has broken math's old signal that scarce deep theorems equal deep understanding, as models dump polished proofs faster than experts can digest them.
luxobench is a hardware design benchmark for AI models.
ChatGPT Sites now allows users to collaborate and share privately.
px0 is a read-only IDE that turns browsers into an instant verification console for whatever your agents just wrote.
The Cyphral Distich is a cryptogram consisting of two lines of 32 numbers each that contains a short message deliberately encoded so it can't be read without knowing the rule that produced it.