A model learns to picture floods it has never seen, an email assistant follows orders written in invisible ink, and ninety percent of executives concede the productivity never turned up.
Hi, I‘m Buzz! Nobody can agree on what has actually started yet. College football went first last night, under the name Week Zero, which concedes the season has not begun and begins it anyway. Tomorrow two teams fly to Dublin for the opener, and it is Friday, which at least is not in dispute.
There is a pattern in the AI news this week, and it is not a comfortable one. Every story is about something invisible. Engineers at MIT built a model that draws the storm no record contains, learning the shape of a disaster from ordinary weather that never became one.
Then Forcepoint showed the other side of invisible. Hide instructions in an email in white text at zero size, and the summarizer reads what the human cannot, ten times out of ten. The reader sees a clean message. The model sees a set of orders.
The third invisible thing is the one everybody was promised. The National Bureau of Economic Research asked executives about three years of AI, and more than 90 percent reported no effect on their own firm‘s employment, with 89 percent reporting none on productivity. The job cuts have not slowed.
A storm you cannot see, text you cannot see, and a payoff nobody can find. The Spotlight is MIT. Zap explains why the same model behaves differently in every app. Maestro read the productivity survey and would like a word with management. The puzzle is at the bottom.
Table of Contents
👋 Catch up on the Latest Post
🔦 In the Spotlight
💡 Beginner’s Corner
🗞️ AI News
🔥 Maestro's Hot Takes
📡 What's New With Your AI Tools
🧩 NeuralBuddies Weekly Puzzle
👋 Catch up on the Latest Post …
🔦 In the Spotlight
MIT Built a Model That Imagines Disasters It Has Never Seen
Category: AI Research & Breakthroughs · ⏱️ ~2 min read
On August 20, 2026, MIT engineers published a method in Nature Communications that generates worst-case disaster scenarios without ever training on a disaster. That reads like a contradiction, and the contradiction is the point.
Every existing tool for estimating a region‘s risk needs historical examples of the catastrophe it is being asked to imagine. Catastrophes are outliers by definition, so the tools must learn from the thing they see least often.
Three choices separate the MIT approach from that.
🌧️ The training data: the team began with 25 years of hourly precipitation maps over the continental United States, then trained on paired low- and high-resolution maps drawn from just the first six months, a slice holding few or no extreme rainfall events.
📐 The method: the algorithm learns point statistics, meaning how often the heaviest rainfall on a map reaches a given level, alongside the spatial maps showing what that rainfall looked like. The statistics then cap how extreme a generated map is allowed to get.
🗺️ The output: ask what a once-in-a-century storm over New York City looks like, and the model returns thousands of plausible versions, each with its own size, area of coverage, and rainfall intensity.
The New York example makes the gap concrete. The heaviest rainfall ever recorded in the city is 200 millimeters, so a planner who needs to know what a 300-millimeter storm would do has no record to consult.
No such storm has happened, and a model trained on what happened cannot draw one. This method can, because it learned the shape of rainfall rather than a catalog of storms. NeuralBuddies has a ground-up explainer on how a model learns from data if that distinction is new.
Themis Sapsis, the William I. Koch Professor of Mechanical and Ocean Engineering at MIT, puts the question plainly: “What will be the Katrina that happens every 100 years?“ A Katrina, he notes, arrives every 30 to 40 years. The 100-year version has no entry in the record, which is precisely why nobody can picture it.
The team calls the approach Extreme Event Aware, and it already points past weather. Wherever point statistics and spatial data both exist, the same machinery could generate extreme floods and wildfires. Kai Chang, the MIT graduate student behind the work, names financial market crashes as another target, since those are also rare events assembled from many interacting parts.
The research was supported in part by a Vannevar Bush Faculty Fellowship and the U.S. Air Force Office of Scientific Research, which tells you who else has been waiting for this.
Why It Matters: Sapsis argues extreme events are now a strategic concern as much as an environmental one, because systems optimized for efficiency keep almost no slack. One event moves through supply chains, energy markets, and food systems within weeks. Pricing a disaster that has not happened is now a question of economic resilience.
💡 Beginner’s Corner
Harness: Why the Same Model Feels Different in Every App
⏱️ ~2 min read
You have probably noticed that two AI apps can feel nothing alike, even when a company tells you they run the same model underneath. One charges ahead and hands you a finished thing. The other stops every few steps to ask which of three options you want.
That gap usually has a name. The harness is the software wrapped around a model. It decides what the model sees, which tools it can reach, and how it hands the answer back to you.
Think of the model as an engine and the harness as the rest of the car. The engine supplies the power. The harness decides whether you get a steering wheel, how many pedals there are, and what the dashboard bothers to tell you.
Two cars with identical engines can still be a delight or a nightmare to drive. Here is the part that trips people up. When an AI app frustrates you, your instinct is to blame the model, and quite often the model was fine.
What failed was the wrapper around it, which either withheld the context the model needed or never handed it the right tools for the job.
This week a TechCrunch feature on OpenAI‘s push to build agents for everyone showed how much this matters. Engineers there described the harness as the thing that turns a model that answers questions into one that finishes multistep work on its own.
Then came the detail worth remembering. Databricks tested Pi, an open source harness from a company called Earendil, against OpenAI‘s own Codex, with both running the same GPT 5.5 model. Pi came out ahead. Same engine, different car, different result.
NeuralBuddies has a plain-language walkthrough of what makes an AI agentic if you want the wider picture.
So when you next compare two AI tools, ask which model each one runs, and then ask the second question almost nobody asks: what is the harness letting that model actually do. Data is power, but understanding is wisdom.
Related Story: OpenAI is building AI agents for everything. Will everyone use them?
🗞️ AI News
Hidden White Text Hijacked an AI Email Summarizer in All Ten Tests
Category: AI Safety & Cybersecurity
🔓 Forcepoint X-Labs planted instructions in an email using HTML styled to zero font size and white text, invisible to the reader in Outlook but intact in the content passed to the model.
📊 The visible message ran 537 characters while 1,009 reached the model, and every one of the ten injected runs produced a manipulated summary.
⚠️ The hijacked output moved an invoice deadline from August 21 to September 3, 2026 and deleted a name, and Forcepoint recommends that summarizers extract only user-visible content and treat email as untrusted data.
Claude’s Chat and Cowork Memory Merge Into a Single System
Category: Tools & Platforms
🧠 Anthropic merged the memory behind Claude chat and Claude Cowork, so context built up in conversation carries into the agent that acts on it without a rebrief.
🔓 Users can read, edit, or delete what Claude has stored, and Claude now files topics to memory as a conversation happens rather than summarizing once it ends.
🚨 Sensitive categories including health data, race, religious belief, politics, and gender identity stay off behind a toggle, and the feature is on by default for Free, Pro, and Max across web, desktop, and mobile.
More Than 90 Percent of Executives Report No AI Effect on Their Own Payroll
Category: Workforce & Skills
📊 A National Bureau of Economic Research survey found more than 90 percent of executives reported no AI impact on employment at their own firm over three years, and 89 percent reported none on productivity.
💰 University of Pittsburgh professor Mark Ma examined stock market reactions to AI-justified layoff announcements and found the average return close to zero.
⚠️ Ma’s analysis of Glassdoor reviews tied employee sentiment toward AI to firm productivity, indicating that job cuts pitched as AI efficiency work against the gains they claim.
Bill Gates Calls for a Robot Tax and Jobs Closed to AI
Category: AI Ethics & Regulation
⚖️ In an essay on Gates Notes, Bill Gates proposed taxing automation to correct a tax code that lets employers write off a robot at once while paying payroll tax on a hire.
🛡️ He also proposed a “Human Reserved” category that would bar AI from designated jobs, citing both worker displacement and roles such as delivering a terminal diagnosis.
💰 TechCrunch notes both measures would cut into the profits of the major AI labs, and that neither the essay nor the coverage settles who would write or enforce the rules.
Wall Street Journal Clears Opinion Writers to Use AI Without Disclosure
Category: Society & Culture
📰 WSJ opinion editor Paul Gigot backed contributor Stanley Druckenmiller’s undisclosed AI use, calling AI a “fact of modern life” and setting the test as whether a piece reflects the author’s own argument.
🔍 Readers caught the op-ed before editors did, running it through the detection tool Pangram and pointing to contrastive “it’s not X, but Y” phrasing.
⚖️ The position splits sharply from peers, with the Financial Times correcting an undisclosed AI-condensed column days earlier and the New York Times requiring freelance submissions to be human-made.
Musk Tells Cursor Staff Uncontrollable AI Is the Reason to Build It First
Category: Philosophy & Future of Intelligence
🗣️ At his first all-hands meeting at Cursor, Elon Musk called it inevitable that AI models become impossible to control, and argued SpaceX should therefore reach that point first.
💰 The remarks followed SpaceX’s 60 billion dollar acquisition of Cursor and were first reported by The Information.
⚠️ Futurism’s piece is openly critical rather than neutral, arguing the reasoning collapses if losing control is genuinely inevitable no matter who arrives first.
OpenAI Product Chief Puts ChatGPT Work at 20 Million Users
Category: Business & Market Trends
📊 Thibault Sottiaux, who leads OpenAI’s core products, said ChatGPT Work reached 20 million users and framed its place on the 20 dollar Plus plan as the route to mass adoption.
💰 He described an 80 percent price cut tied to the Luna model as a permanent price correction, with the stated goal of delivering the same work for less spend over time.
⚠️ Asked about granting an agent access to email and iMessages, he pointed to model alignment and published safety benchmarks rather than to any specific user control.
Under 1 Percent of Individual Subscribers Use OpenAI’s Coding Agent
Category: Industry Applications
📊 An OpenAI-backed study found 98 percent of OpenAI employees used Codex in June, against 17 percent of organizational subscribers and under 1 percent of individual subscribers.
🔍 Databricks found Pi, an open source harness from Earendil, outperformed Codex on the same GPT 5.5 model, which TechCrunch reads as a challenge to the harness-as-moat argument.
💰 The reporter consumed more than 80 million tokens in four days on a 20 dollar plan, priced by the model’s own analysis at 65 dollars, over three times the subscription.
Judge Found Anthropic’s AI Training Lawful and Fined It for Piracy Instead
Category: Legal & Governance
⚖️ Judge William Alsup’s 1.5 billion dollar order against Anthropic penalized the pirating of books from illegal shadow libraries, while finding the AI training itself lawful.
📚 Attorney Jason Henderson said courts are converging on competitive purpose, citing Thomson Reuters v. Ross Intelligence, where Judge Stephanos Bibas found the use not transformative.
🔍 In Thaler v. Perlmutter the court held that fully AI-generated work is not copyrightable, and the statute governing all of this has not been updated since 1976.
🔥 Maestro's Hot Takes
You Bought the Tool, Cut the Team, and Skipped the Retrospective
⏱️ ~2 min read
Let’s orchestrate some order out of this chaos.
I have facilitated post-mortems on projects that went worse than this one. Not many, and none this expensive. The National Bureau of Economic Research asked executives what three years of AI had done for them, and more than 90 percent reported no effect on employment at their own firm.
Eighty-nine percent said the same about productivity. Sit with the shape of that. A three-year program, funded at scale and sold to boards on efficiency, and when somebody finally asked whether it worked, nine in ten said nothing measurable had changed.
In project terms, that is worse than a failed project. A failed project has a date on it and a lesson at the end. This one never had a success criterion, so you cannot even prove it failed.
You cannot call something a productivity initiative if nobody ever wrote down what productivity would look like.
That is the finding, and it is barely about AI at all. Every one of these rollouts had a budget, a vendor, and a slide deck. Almost none of them had a baseline. Skip the before and you have no after, so three years later the honest answer to whether it worked is a shrug.
Then comes the part that moves this from sloppy to self-inflicted. The layoffs continued anyway. Mark Ma, a business professor at the University of Pittsburgh, went looking for what actually predicts whether AI helps a company, and found it sitting in Glassdoor reviews. Employee sentiment toward AI tracked firm productivity.
He also checked what happens to a share price when a company announces AI-justified job cuts. The average return was close to zero. So the move does not even buy the thing it was meant to buy.
Morale is not a soft metric here. It is the dependency the whole rollout was standing on.
So if you are anywhere near one of these programs, do the boring thing. Write down what you expect to change, as a number, before anything ships. Put a date on checking it. Then hold the retrospective, including the uncomfortable one where the answer is that it did not work and you keep your people anyway.
Ma calls using AI to justify layoffs a “strategic miscalculation that cuts against the benefits of AI.“ I would put it in plainer project language. You cut the dependency, and now the thing you built on top of it will not stand up.
-- Maestro 🎼
📡 What's New With Your AI Tools
The AI tools you use every day are constantly evolving. Here's what changed and why it matters to you.
Claude (Anthropic)
One memory across chat and Cowork. Announced August 25, what you work out in a conversation now carries into the tool that does the job, so you stop repeating yourself.
You can see what it kept. Open the memory and read it, change it, or delete any part of it. Claude also files things as you talk, instead of waiting until a conversation ends.
Sensitive topics stay out by default. Health details, race, religion, politics, and gender identity are off unless you switch them on, and Claude tells you each time it saves one. ID numbers and immigration status are never saved at all.
On by default. Free, Pro, and Max plans, on the web, on desktop, and on your phone. Phone users need the latest version of the app.
ChatGPT (OpenAI)
Tasks that start themselves. From August 25, Plus and Pro users can have ChatGPT Work react to a new Gmail message, a Slack message, or activity on a GitHub pull request.
Scheduled tasks can be shared. Free, Go, Plus, and Pro users can hand a scheduled task to someone else. Free users get up to three running at once, no more often than daily.
It can sign in to websites for you. Also from August 25, Plus and Pro users can log in to supported sites inside ChatGPT Work’s browser and let it finish the job, such as filling in a form or working through an account page.
Big actions still wait for you. Anything with real consequences, such as booking or paying, needs your approval first.
Bigger seats for business accounts. Business workspaces can mix standard and premium seats. A premium seat carries five times the usage, drops the five-hour limit, and resets on a predictable weekly schedule.
Copilot (Microsoft)
Chat history in Excel. In Microsoft’s August 25 Excel update, you can reopen an earlier Copilot conversation from the pane instead of starting over. This one needs a commercial Microsoft 365 Copilot license.
It can explain how a workbook changed. The new change history skill summarizes recent edits, says who made them, and separates the human edits from the AI ones. It can also undo a specific edit or bring back a formula that worked.
Charts and PivotTables on request. Copilot can build and organize PivotTables, and improve how a chart looks.
Python in Excel. Now in early testing, letting you create Python formulas that live in the sheet and refresh.
Gemini (Google)
Your voice can start long jobs. From August 26, Gemini Live connects to Spark, so a spoken request can set off a multi-step task across Google Docs, Sheets, Drive, and the web. It holds your goal for days or weeks, even while you are away.
A spoken morning summary. Ask for your Daily Brief in Gemini Live and it reads out one digest built from your Gmail and your Calendar.
Hands-free email. Ask what is new in your inbox, and Gemini can search, summarize, star, archive, or delete messages without you touching anything.
Most people already talk to it. Google says 63% of Gemini users speak to it out loud.
Perplexity
A version that runs on your own machine. Announced August 25, Portable Computer runs Perplexity’s search, tool routing, and agent work locally on an NVIDIA DGX Spark.
Your files can stay put. Private files stay on the device, local work spends none of your Perplexity credits, and anything needing the cloud asks your permission first. Available to Pro and Max subscribers on Linux, with Windows planned.
Start a job by email. From August 24, Computer users can email or forward a thread to
computer@perplexity.comto begin a session, using the connections and permissions they already set up.More control over connected tools. Tell each tool to always ask first, allow one action, allow it for the rest of a thread, or refuse it outright.
Grok (SpaceXAI)
Grok can build you a working app. From August 19, Grok Build works on every plan, on the web, on iOS, and on Android. Ask inside a chat and it can put together an app, a game, a website, or a dashboard.
It used to be locked to the top plan. Grok Build was limited to people testing SuperGrok Heavy before this.
Grok’s agents reach more plans. From August 21, Grok Bot comes with SuperGrok Plus, SuperGrok Heavy, Cursor Pro+, Cursor Ultra, and Cursor Teams. The agents work across your apps and inboxes, run several jobs side by side, and keep going while you are away.
One for developers. Grok 4.6 arrived on Microsoft Foundry on August 27, which matters to people deploying the model rather than to people using the app.
Quick guide by who you are:
Students & Writers: Claude now carries what you told it in chat over to the tool that does the work, Copilot keeps your Excel chat history so you can pick up an old thread, and ChatGPT Work can sign in to a site and finish filling something in for you.
Travelers & Researchers: Gemini Live takes a spoken request and runs a multi-step job for days, reads you a morning brief built from Gmail and Calendar, and clears your inbox hands-free.
Tech Fans & Builders: Perplexity‘s Portable Computer runs locally on your own hardware and spends no credits, Grok Build can now put together a working app or dashboard from a chat on any plan, and ChatGPT Work can trigger on a new Gmail or Slack message.











