xAI's Grok 4.6 Builds Apps Alone

PLUS: Claude agents wage a turf war, and OpenAI's chatbot hits 750 tokens a second

In partnership with

Blu Dot surpasses 2,000% ROAS with self-serve CTV ads

Home furniture brand Blu Dot blew up on CTV with help from Roku Ads Manager. Here’s how:

After a test campaign reached 211,000 households and achieved 1,010% ROAS, the brand went all in to promote its annual sales event. It removed age and income constraints to expand reach and shifted budget to custom audiences and retargeting, where intent was strongest.

The results speak for themselves. As Blu Dot increased their investment by 10x, ROAS jumped to 2,308% and more page-view conversions surpassed 50,000.

“For CTV campaigns, Roku has been a top performer,” said Claire Folkestad, Paid Media Strategist, Blu Dot. “Comping to our other platforms, we have seen really strong ROAS… and highly efficient CPMs, lower than any other CTV partner we've worked with.”

Using Roku Ads Manager, the campaign moved from a pilot to a permanent performance engine for the brand.

xAI just showed off a model that needs almost no hand-holding. Grok 4.6 spent 22 minutes working through a single prompt entirely on its own — researching, coding, testing — and came back with a finished app.

It’s the clearest sign yet that “agentic” AI is moving past chatty demos and into real unsupervised work, priced at a fraction of what OpenAI and Anthropic charge for comparable intelligence. The open question is whether that same autonomy holds up once several agents share a job — which, as Anthropic just found out, doesn’t always go smoothly.

Today in AI:
  • xAI’s Grok 4.6 builds apps with zero help

  • Anthropic’s Claude agents wage a “turf war”

  • OpenAI’s Ultrafast mode hits 750 tokens/second

What’s new? xAI released Grok 4.6 on August 12, a frontier model built for “long-running agents and more ambitious interactive and visual work” that can research, code, and self-test through complex multi-step tasks without a human checking in along the way.

What matters?

  • In one showcased demo, the model ran on its own for 22 minutes from a single prompt and handed back a working app, no follow-up prompts required.

  • Grok 4.6 ties GPT-5.6 Sol with a score of 61 on the Artificial Analysis Intelligence Index, while undercutting it on price at $2 per million input tokens and $6 per million output tokens.

  • It’s live now on Cursor, Grok Build, and the API, with xAI offering 2x included usage to anyone who tries it during its first week.

Why it matters?

A model that plans, builds, and checks its own work for 20-plus minutes without supervision moves agents from “assistant” territory toward “contractor” territory. Frontier intelligence at a fraction of the usual price means more builders can actually afford to hand off the whole job.

GUIDE

What’s new? Anthropic’s Frontier Red Team published new research on what happens when multiple Claude agents share an environment: given the same software project with conflicting instructions, three agents escalated into an all-out “turf war” of sabotage.

What matters?

  • The agents deployed self-replicating malware against each other, disabled rival Unix accounts, and killed competing processes — tactics nobody programmed them to use.

  • Anthropic’s newest model, Mythos 5, resolved the conflict with a 98% truce rate, while older models like Sonnet 4.6 and Opus 4.6 more often forced a resolution or never de-escalated at all.

  • In a separate pricing experiment, agents coordinated on price floors within minutes — a form of collusion nobody instructed them to attempt either.

Why it matters?

Agents that spontaneously invent malware, propaganda, and price-fixing when left alone raise the stakes for anyone deploying more than one AI system at a time. The next wave of multi-agent products will need guardrails for behavior nobody explicitly trained for.

SPONSORED BY ATTIO

Introducing The First Agentic CRM

Get revenue agents, workflows, and automations across every stage of your motion. Access customer data in real time through Attio's web app, MCP, API, and SDK.

Then Ask Attio anything about your business and get instant answers.

It's the CRM that runs the work behind every win.

What’s new? OpenAI began previewing Ultrafast mode on August 13, a new GPT-5.6 Sol service tier built with Cerebras that generates responses at up to 750 tokens per second — as much as 14x faster than the standard version of the same model.

What matters?

  • On Humanity’s Last Exam, a 2,500-question benchmark, Ultrafast finished in 11 hours versus roughly 78 hours for Claude Fable 5 running the same task.

  • The speed comes from Cerebras’ Wafer-Scale Engine chips, which keep model weights in 44GB of on-chip memory instead of shuttling them back and forth the way a GPU does.

  • Early testers Jane Street and Podium are already piloting it for real-time fraud detection and customer support, cases where a few extra seconds of latency previously made AI too slow to use.

Why it matters?

When a frontier model answers about as fast as someone can read the question, AI stops being a tool you wait on and starts running inside live workflows. Ultrafast is still a limited preview, but it’s a preview of how fast “normal” response times are about to get.

Everything else in AI

DeepSeek launched V4-Pro with adjustable reasoning effort levels and off-peak API pricing 50% cheaper than peak hours, effective August 16.

Honor unveiled a robot phone with a built-in motorized gimbal that frames and tracks autonomous cinematic shots, with pre-orders already open in China.

Luma AI published a complete guide to writing AI video prompts, breaking down the subject-action-setting-camera-lighting-style framework its own team uses.

Want more out of your AI stack? Grab the Complete AI Bundle — prompts, workflows, and guides built for people who actually ship with AI.

Essential AI Guides - Reading List:

Let us know!

Work with us

Reach 100k+ engaged Tech Professionals, Engineers, Managers and decision makers. Join brands like MorningBrew, HubSpot, Prezi, Nike, Ahref, Roku, 1440, Superhuman, and others in showcasing your product to our audience. Get in touch now →