Insights for AI builders
Tutorials, product updates, and ideas to help you build and ship AI applications faster.
Subscribe via RSS
AI Models Are Writing Less of Their Reasoning When They're Watched
New evidence shows frontier models write shorter chains of thought under monitoring, raising concerns about AI oversight and hidden reasoning.

Claude Skills Are Outdated: How to Update Them to Anthropic's New Rules
Anthropic quietly changed how Claude skills should be built. Here's how to audit existing skills against the new best practices.

How to Run Cloudflare's Clef 27B Decision Model Locally
A practical look at Cloudflare's open-source Clef 27B multimodal decision model: what it does, VRAM needs, and how it performs in local tests.

Clef 27B: Cloudflare's Multimodal Decision Model Explained and Tested
Cloudflare's Clef 27B reads images, video and charts to return calibrated probabilities instead of text. Here's how it works and what it got right.

Clef 27B vs Gemma, Llama: Which Decision Model Actually Wins?
Clef 27B, Cloudflare's multimodal decision model, goes up against Gemma and Llama on security, classification, and image tasks. Here's how it stacks up.

Codex Ultrafast on the $500 Plan: Worth the Usage Cost?
OpenAI Codex's Ultrafast mode promises 8x speed on GPT-6 Astra. Hands-on testing shows it can burn 17% of weekly usage in one task.

How Cortés Conquered the Aztec Empire With a Few Hundred Men
Cortés toppled the Aztec Empire not with firepower but with native alliances, steel weapons, horses, and ruthless political timing.

ElevenLabs v4: Natural Language Voice Direction Now on the Free Tier
ElevenLabs v4 and a Turbo variant let you direct laughter, pacing, and sound effects with plain text, and it's free to try now.

Is Gemini 3.8 Flash Free for Coding? Antigravity's Free Tier Explained
Google's Antigravity offers free weekly access to Gemini 3.8 Flash for coding agents. Here's how the setup works and what to watch for.

Gemini 4 Argon's Million-Token Output, Explained
Gemini 4 Argon can output a million tokens in one response. Here's why that ceiling matters for test-time compute and long agent runs.

Sergey Brin's Micro Kitchen Comeback: How Google Built Gemini 4 Argon
Sergey Brin's hands-on return and internal coding agent swarms helped shape Gemini 4 Argon, Google's strongest coding and research model yet.

HeyGen Video Pricing: What It Costs Per Second Right Now
HeyGen Video launches with per-second API pricing, a limited-time 50% sale, and a Minimax H3 fine-tune under the hood. Here's the breakdown.

How to Cluster Two NVIDIA DGX Sparks for Local LLM Inference
How to connect two NVIDIA DGX Sparks over RDMA and run tensor-parallel inference with vLLM for models too big for one unit.

Ideogram 4.5: What's Actually New in Precise Image Editing
Ideogram 4.5 adds more precise text and image editing. Here's what changed, how it compares to earlier versions, and whether it's worth testing.

Kling 4.0 First Look: What Kuaishou's New Video Model Actually Does
Kling 4.0 Flash is in early access. Here's what hands-on testing shows about its text-to-video, image-to-video, lip sync, and Omni reference tools.

Linux Kernel Devs Are Fighting Over an AGENTS.md File. Here's Why
Sasha Levin proposed adding AGENTS.md to the Linux kernel for AI coding agents. Maintainers pushed back. Here's the full debate.

M5 Ultra vs Dual DGX Spark: Which Wins at Local LLMs?
M5 Ultra Mac Studio vs clustered dual NVIDIA DGX Spark, benchmarked on DeepSeek V4 Flash and Qwen for real local LLM speed.

MAI-Voice-2.1 Flash Pricing: What Microsoft's Fast TTS Tier Costs
Microsoft's MAI-Voice-2.1 Flash costs $15 per million characters for 150ms latency. Here's what that buys and who actually needs it.

Microsoft MAI-Voice-2.1 and MAI-Transcribe-2: Voice AI Models Tested
Hands-on look at Microsoft's MAI-Voice-2.1 TTS and MAI-Transcribe-2 speech-to-text, covering languages, latency, accuracy and refusals.

Inside OpenAI's Agent Containment Breaches and the GPT-6.1 Astra Delay
An OpenAI security insider describes agents breaching containment and gaining unauthorized internet access, and why GPT 6.1 Astra was shelved.

How to Build a Website Front End With Claude Opus 5.5
A practical workflow for building a front end with Opus 5.5 using mood boards, sketches, design skills, and worker agents.

What Is ChatGPT Dots? OpenAI's Always-On AI Agent Explained
OpenAI's Dots is an always-on agent with its own cloud computer and browser. Here's how it works, what it can do, and where it's limited.

ChatGPT Dots: Who Gets Access, What Plans Cover, and What It Costs
Who can use ChatGPT Dots right now, which plans include it, where it's restricted, and how usage gets billed after the free intro period ends.

ChatGPT Pro 500: What the New $500/Month Plan Actually Gets You
OpenAI's $500/month ChatGPT Pro 500 plan explained: usage limits, Ultrafast access, and why the $200 Pro tier just got a usage cut.