📡 TLDR Newsletter

← Zurück
Qwen3.8 Is Going Open-Weight (1 minute read)

Alibaba announced Qwen3.8, a 2.4-trillion-parameter model slated for an open-weight release. A preview version is available through Alibaba's Token Plan, Qoder, and QoderWork.

Apple Sends Legal Letters to Dozens of OpenAI Employees (2 minute read)

Apple has reportedly sent legal letters to dozens of former Apple employees now working at OpenAI. The letters instruct the employees to preserve potentially relevant documents and communications rela...

How Netflix Built Its LLM Serving Stack (18 minute read)

Netflix details how it deployed and operated LLM inference within its existing production infrastructure. The article covers engine selection, model packaging, API design, deployment strategy, output ...

Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal Help? (6 minute read)

Claude Code and Codex both have a /goal option, but their implementations are fundamentally different. Claude Code implements /goal as a session-scoped Stop hook, so it can catch an early exit, but it...

Demis Hassabis and Google's AI Commitments (18 minute read)

This post examines Demis Hassabis's framework for frontier AI alongside criticism of Google's military agreements. It also covers Alex Turner's resignation after unsuccessfully opposing broad governme...

CI was built for humans. Chunk is built for agents (Sponsor)

Chunk Sidecars by CircleCI catches failures in a 27-second microbuild, before your agent ever reaches CI. 5x fewer tokens. 78% faster than a full pipeline run. Free, agent-agnostic, works with Claude ...

Diffusing Blame: Dale-Constrained Learning Without Weight Transport (3 minute read)

Sakana AI's "Diffusing Blame" is a neural network technique that enforces Dale's principle while enabling effective learning in image classification and reinforcement learning tasks. Their approach, b...

Kimi Code CLI (GitHub Repo)

Kimi Code CLI is an AI coding agent that runs in the terminal. It can read and edit code, run shell commands, search files, fetch web pages, and choose the next step based on feedback. The CLI works w...

Claude Fable 5 will be included in all Max and Team Premium plans (1 minute read)

Claude Fable 5 will be included in all Max and Team Premium plans at 50% of limits starting July 20. Pro and Team Standard users will continue to have access to Fable via usage credits. They will rece...

Alibaba open-sources its AI chip software stack at WAIC, targeting Nvidia's CUDA lock-in (3 minute read)

Alibaba is open sourcing the software stack for its Zhenwu series of AI chips. The move will lower migration barriers for developers currently locked into Nvidia's CUDA ecosystem. If China wants AI in...

How to manage AI investments in the agentic era (5 minute read)

OpenAI highlights that falling AI costs should be evaluated through useful work per dollar, not token prices alone, emphasizing visibility, outcome-based ROI, governance, scalable workflows, and capac...

A Chinese AI startup is about to hit $1bn in sales while giving its best models away for free (3 minute read)

A large share of Z.ai's income comes from on-premises deployments for state-owned enterprises and financial institutions, as well as its fast-growing cloud business.

Why the first GPU financiers are turning to inference chips in a $400 million deal (4 minute read)

General Compute, an AI inference cloud startup, is using inference-specific chips as collateral for a $400 million loan.

GitLab 19.2 release notes (6 minute read)

GitLab 19.2 introduces general availability for GitLab Duo CLI, custom AI flows, scheduled pipeline execution policies, CI Expert Agent, and fine-grained personal access tokens, alongside new AI-power...

I burned all my tokens researching how to save tokens (12 minute read)

Use cheap models to find information, accurate models to verify it, and then use deep research last.

Reviewing AI-generated code (9 minute read)

AI-generated code shifts code review away from syntax and implementation details toward validating requirements, architecture, security, and correctness. Effective reviews focus on whether the code so...

Isolation levels: Read Committed, Repeatable Read, and Serializable explained (6 minute read)

Database isolation levels make different tradeoffs between concurrency and consistency by controlling which anomalies concurrent transactions are allowed to observe. Read Committed prioritizes through...

Moonshot AI Plans Hong Kong IPO After Kimi K3 Model Debut (3 minute read)

Moonshot AI plans to list on the Hong Kong Stock Exchange within six months.

Kimi K3 has received far more love than expected (1 minute read)

Kimi is temporarily pausing new subscriptions and prioritizing compute for current members.

Google preparing Skills and Gemini Live for web rollout (2 minute read)

Google plans to introduce Gemini Live to desktop and web, extending real-time voice capabilities beyond mobile.

Open source just reached frontier code review 🔍 (Sponsor)

PR-AF places #2 of 42 on Martian's Code-Review-Bench — ahead of CodeRabbit, Copilot, and Devin. It's a self-hosted, drop-in GitHub Action: the harness plans a review per PR, runs reviewer agents in pa...

Securing the AI supply chain on GKE: Introducing k8s-aibom for automated AI BOMs (6 minute read)

k8s-aibom is an open-source Kubernetes controller that detects running AI workloads and generates CycloneDX 1.6 ML-BOMs without privileged access, sidecars, or developer changes. It provides runtime v...

Wigolo (GitHub Repo)

Wigolo, a new open-source tool, lets AI agents search, crawl, and extract web data entirely locally without API keys or usage fees, running as an MCP server that integrates with Claude, Cursor, and ot...

Answer any cost question faster with the Cloud Cost skill in Bits Chat (5 minute read)

Datadog launched a Cloud Cost skill in its Bits Chat AI assistant that lets FinOps and engineering teams investigate cloud, AI, and SaaS spending through conversational queries, with the ability to an...

What happens in the milliseconds after you tap pay (8 minute read)

Databricks demonstrates a fraud detection app that scores credit card transactions in 27 milliseconds median latency by combining Model Serving route optimization with Lakebase Postgres for real-time ...

AWS Lambda announces self-managed code storage (2 minute read)

AWS Lambda now supports self-managed Amazon S3 code storage, allowing functions and layers to reference deployment packages directly from customer-owned buckets instead of Lambda-managed copies.

Amazon EC2 now surfaces the public SSM parameters associated with public AMIs (2 minute read)

Amazon EC2 now includes associated AWS Systems Manager Parameter Store parameters in public AMI metadata, simplifying discovery and configuration updates.

Cloudflare WAF protects WordPress applications from two high-severity vulnerabilities (3 minute read)

Cloudflare deployed Web Application Firewall protections on July 17 for two critical WordPress vulnerabilities—an unauthenticated Remote Code Execution flaw and a SQL Injection vulnerability—affecting...

Reach the people who make tech buying decisions (Sponsor)

One of the biggest challenges in digital advertising is reaching the people who actually make buying decisions. Founders, CTOs, engineering leaders, and product managers read TLDR every day to stay on...

Meta in Talks to Lease Computing Power to Anthropic in Potential $10 Billion Deal (5 minute read)

Anthropic proposed a computing deal with Meta in June that could be worth as much as $10 billion over two years. Meta is considering the deal, which would involve monthly payments and the option to op...

SpaceX in Talks to Provide Computing Power for Pentagon's AI Push (3 minute read)

SpaceX is in talks to provide the Defense Department with computing capacity at a cost of up to several billion dollars. Some national security officials have raised concerns that the Pentagon is beco...

India's first privately-developed rocket reaches orbit on dramatic debut launch (12 minute read)

Skyroot Aerospace's Vikram-1 rocket, India's first fully commercial satellite launcher, launched into a 280-mile-high orbit from an island spaceport in the Bay of Bengal on Saturday. The launch was de...

China Joins Rush to Rethink the Smartphone for the AI Era (5 minute read)

Several companies are now racing to develop devices with built-in AI agents. ZTE Corp. has unveiled a lineup of co-designed smartphones with built-in AI services. The NaviX Ultra, which it calls the w...

Tune it, deploy it, own it: Crusoe Serverless Fine-Tuning is now live (Sponsor)

Fine-tune leading open models (Qwen, DeepSeek, Gemma, gpt-oss) on your proprietary data with no clusters to provision and no surprise bills. Crusoe's token-based pricing ends the moment your model sto...

How Anthropic runs large-scale code migrations with Claude Code (14 minute read)

Code migrations were typically multi-year endeavors until recently. Claude Code's capabilities change the math for these types of projects. This post details the six-step process Anthropic uses to per...

scroll-world (GitHub Repo)

scroll-world is a skill that turns any brand into a scrollable 3D world. It works with Claude Code, Codex, and any SKILL.md-compatible agent. The skill builds immersive, control-scrubbed 'fly through ...

Maybe Intelligence Ain't All That (4 minute read)

AI chatbots are now smarter and more capable than a lot of people, but there has yet to be an explosion or take-off in usage. Coming up with ideas was never the hard part. Most problems can't be worke...

These AI-Native Companies Have Tiny Staffs and Fewer Bosses (7 minute read)

The newest generation of AI-infused companies offer a vision of how work could soon be structured in other American corporations. They have fewer workers, more on-staff engineers, and a flatter struct...

Qwen3.8 has launched (1 minute read)

The massive 2.4T parameter model is also going open weight soon.

Apple, Nvidia vie for title of world's most valuable company (2 minute read)

Apple briefly became the world's most valuable company on Friday.

An AWS billing bug sent users estimated charges of up to $2.5 trillion (2 minute read)

Amazon has confirmed a unit pricing error and is recomputing all estimates.

Are the LLM Wars the Database Wars? (2 minute read)

The winners of the database wars weren't the ones everyone was arguing about.

Domain-specific harnesses (2 minute read)

Different workloads perform better with certain combinations of models and harnesses.

AI Mania Is Eviscerating Global Decisionmaking (38 minute read)

The people in charge either have no plan or see no path forward other than keeping their heads down.

An insider look at Atlassian's AI-native vision for service management (Sponsor)

Bolting AI onto legacy playbooks will rarely deliver the expected results. What's needed is a reimagining service of management from the ground up.Grab a copy of Atlassian's Shatter the service quowhi...

NVIDIA OpenShell Secures the Agent. Who Governs the Fleet? (8 minute read)

NVIDIA OpenShell is an open-source runtime that sandboxes AI agents with kernel-level policies controlling file access, process spawning, and network traffic, validating the principle that agent secur...

Introducing Apache Spark 4.2 (7 minute read)

Apache Spark 4.2, now available in Databricks Runtime 19 Beta, introduces governed metric views as a native semantic layer, vector similarity search primitives for AI retrieval, and first-class change...

One voice. Every device you own. (Sponsor)

Wispr Flow works the same on your Mac, your PC, and your phone. Start an email on your laptop, finish a Slack reply on your phone. Same hotkey, same clean output, 4x faster than typing.89% sent with z...

Turso is building a modern version of Postgres in Rust (6 minute read)

Turso Limbo is a new PostgreSQL-compatible database written in Rust that is designed to modernize PostgreSQL's architecture while maintaining wire-protocol compatibility and existing tooling. The proj...

How to Encrypt Terraform State Files (13 minute read)

Terraform state encryption protects sensitive infrastructure data by securing state files in transit with TLS and at rest through backend encryption, while OpenTofu adds built-in state and plan file e...