This is Robin Sloan’s lab notebook. It’s about media and technology, creative computing, AI aesthetics, & more. Here's the RSS feed. My email address: robin@robinsloan.com
The OpenAI-Hugging Face incident remains THE fascinating event of the summer, maybe the year; deeper investigation has revealed its rich, strange structure.
But, notice:
Over the course of this investigation, OpenAI provided us with the dump of ~1.2 million entries from the main message board and the dataset of ~1300 transcripts we describe below, as well as free API credits for GPT-5.6 Sol for analysis.
How do you make sense of ~1.2 million agent messages and ~1300 very long LLM agent activity transcripts? With another LLM, of course.
This is a pattern that recurs in this domain. Assembling training data at the scale required by 2020s-era models, no researcher can “read it all”. So, you either (1) don’t bother, or (2) use another LLM to review and filter the data. You can, in principle, use other kinds of models — simpler classifiers — but, increasingly, the kinds of judgments you need to make require the richness of an LLM.
Anthropic’s Insights tool, likewise, uses Claude to read and categorize millions (billions?) of transcripts of people’s interactions with Claude. In addition to making this huge heap of data legible at all, the “LLM in the middle” acts as a privacy buffer: researchers read only Claude-generated summaries, not the original interactions.
I’ve come to think of this as “using tongs”, in the sense of a tool that allows you to manipulate material that you otherwise couldn’t.
Or maybe the better analogy is one of those laboratory gloveboxes, and the boundary being maintained isn’t about atmosphere, but rather scale. Imagine the scientist’s hands ballooning up in size, a million times, as they reach into the chamber:
Containment
It makes me think also of the pantograph, a once-ubiquitous analog tool for changing the scale of a drawing, or any kind of mechanical operation:
When an LLM acts as a “digital pantograph” for text, it can “scale up”—expand a one-sentence prompt into thousands of lines of code — or “scale down”—categorize and summarize millions of messages.
But a real pantograph is a simple, predictable, inspectable tool … and an LLM is nearly the opposite. Notice the risk: a truly sneaky model, asked to scour the transcripts of its cousins for misdeeds, could easily refuse to snitch: “Yep, I read all 1.2 million messages … nothing to see here!”
Even without collusion, you settle for coarse and inflexible analysis. When you tell an LLM to read a bunch of documents and answer questions about them, you get: answers to those questions. When you read a bunch of documents yourself, you also get: new questions! In an investigative mode, this is really important.
I semi-jokingly called our efforts a “slop-vestigation” because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data — over a thousand extremely long transcripts from agents that ran for multiple days — made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn’t mean these agents could be easily used to oversee and understand the incident.
Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them.
We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation.
Anyway, it’s all very weird, and this problem of “how do you make sense of millions of messages (or more) written by AI agents?” is only going to become more widespread and, in many cases, more urgent. It would be interesting to think about LLMs that are kinda dumb, but have ULTRALONG context windows and the ability to make simple judgments across them. What kind of machine would be required to literally “look at the entire OpenAI-Hugging Face incident at once”—hold it all in its head? (Maybe I should peruse this work … )
This conundrum reminds me also of “distant reading”, Franco Moretti’s research program from the 2000s, which was pursued with much cruder computational tools. I wonder if there might be some usueful nuggets waiting in that early work.
A near-ubiquitous reference, or aspiration, for AI agents is J.A.R.V.I.S., the supercapable computer helper in the Iron Man movies. “Finally building my own JARVIS!” is a popular refrain among enthusiastic AI tinkerers. (Not least because it’s flattering: “If my computer is Jarvis, I guess that makes me Tony Stark … ”)
However, the most common application of a personal Jarvis seems to be … tinkering with one’s personal Jarvis. “Gotta get my tools just right” isn’t a new phenomenon, of course, but/and it’s useful to notice its recurrence here.
So I just want to say: Jarvis wants to help you DO something! Or make something, or investigate something.
Continuing a happy rewatch of Star Trek: The Next Generation, I’ve been struck by the prescience and appeal of the crew’s collaborative relationship with their ship’s great anonymous Computer. As cool as Jarvis is, I’ll argue that this bravura scene from TNG is actually a much richer picture of what might be possible with computer collaborators. Geordi solves a mystery using the Computer and the holodeck together — go watch, it’s great.
Steve Krouse, talking about Val Town’s growth in July, writes:
The majority of those customers were referred to us by AI.
Previously, we heard about Buttondown’s first glimmers of AI referral business. Val Town makes even more sense, because it is literally THE place to go to deploy small apps — e.g., the one you just asked Claude to build.
Surely people are already pitching a new cousin to SEO aimed at the chatbots and agents. Beyond the obvious/sensible — “Make your platform legible to language models, with rich clear documentation in Markdown, etc.”—I wonder what kind of (possibly cynical) strategies are being deployed?
An LLM’s “sense of the world” does seem significantly more difficult to game than Google’s search index, but then again, I don’t have that black hat mind … I’m sure there are tricks emerging even now.
The goal of course is to become “the default pick” for X, where X is “a database”, “an email provider”, “an olive oil subscription”, whatever: the one that Claude Code recommends before the user even asks! What a flood of referrals … can you even imagine? I wonder about those default picks for X even now, across models. I’ve seen some broad discussion, but one could just run a set of “which X should I use … ?” questions through OpenRouter to ten different models and see what emerges. Maybe I will do this …
I’m on the road, so just posting a quick link: here is the first episode of KQED’s Dream Machines, the podcast miniseries about AI hosted by me and Alexis Madrigal! Our guest is Annalee Newitz, speaking here as perhaps the world’s foremost imaginer of other minds, and also a great, iconic San Franciscan.
I am both a friend and superfan of Alexis, so the opportunity to collaborate with him on this project has been a delight. Our process: take the content of our text threads out of the Messages app, into the podcast studio. And I’m very proud of the fact that this is a KQED production — which is to say, public conversation, not industry capture. Truly independent.
I’ll have more to say about the project and its intentions, possibly in my next newsletter. KQED will release several more episode over the course of this month — I’ll post links as they arrive.
How delightful to follow a link to the new RSS reader and article archive called Cove, only to discover, right there on the front page, peeking out of the big beautiful demo screenshot …
I spent a day testing prime-agent and ended it with an unpleasant surprise.
The agent automatically discovered my OpenRouter and OpenAI API keys and started using them instead of my OpenAI/Anthropic subscriptions. What made it worse: I couldn’t find any proper way to remove or disable the auto-detected providers and models.
A harness this flexible really needs a kill switch for exactly this scenario. The list of providers, models, reasoning efforts and their settings should be explicitly defined by the user — opt-in, not auto-discovered. Otherwise the whole thing becomes uncontrollable, and potentially expensive.
The image of a computer program as an unruly guest: the minute you leave, they’re rifling the drawers. 2026!!
Apropos of David Bushell’s post, I thought I’d just mention, I still use Sublime Text for all of my programming AND all of my newsletter-ing — I write them as Markdown files first, then send rendered HTML to Buttondown via its great API.
I am, indeed, typing this short post into Sublime Text.
The app is simple and superfast; I have it set up exactly the way I like it (with that setup synced between computers, via Dropbox), and it’s difficult for me to imagine ever switching to anything else.
Here is a piece of software as sturdy and obedient as a cast-iron pan.
[Wall Street Journal reporters] are trained to find the needle in a haystack, but doing so on a breaking news timeline can be challenging. To speed up the document review, the reporters leaned on a pre-built internal tool called WSJPT (a play on ChatGPT). The tool standardizes basic LLM requests across reporting projects, including prompts for summarization, classification, and image description. In this case, the reporters used the tool to summarize every page of every document scraped from the county portal.
“WSJPT” is indeed very cute!
The leverage these tools provide — the absolute ease with which they will dance through dumpsters full of documents — is breathtaking. Of course, this kind of search should only be a starting point … but the point is, previously, this kind of search simply was not possible.
That’s via Context Window, a newsletter about AI’s impact on media and publishing — a new favorite.
As you probably heard, a bullet point recently appeared on the timeline of computers, AI, and maybe everything: AI agents running in a OpenAI’s training environment broke out and hacked the servers of another tech company.
While I understand that architecting and managing these systems is anything but easy, the fact that this was even possible seems CRAZY to me. If research scope and speed are at odds with “my agents have been planning and executing operations on the open internet for weeks, without my knowledge”, then research scope and speed need to change immediately — and it sounds maybe like they have.
Seriously, do watch the video, and, as you do, conjure the creepy recognition that, a few weeks ago, these agents were out there, doing this work, communicating through subtle channels, and nobody knew, not even their operators.
And so, the obvious question arises: what agent swarm is out there working NOW without anyone’s knowledge … and what is it doing?