
Two things AI could break · ↗ garymarcus.substack.com
Scientific collaboration and the world
Gary Marcus had a good piece on “two dire warnings” about AI. The first comes from Terence Tao regarding the recent controversy over the solution to the Navier-Stokes Millennium problem. Tao warns:
We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field.
This controversy over credit—and what OpenAI’s models may or may not have learned from the Codex logs of the mathematicians working on Navier-Stokes—may cause companies that rely on IP protection, such as drug developers, to demand stronger guarantees that model providers will not use their prompts to compete with them. Or it might cause them to go all-in on local models and custom harnesses.
Either way, scientific/mathematical inquiry transitions from being relatively high trust to being low trust and insular, fueled by paranoia about second-rate (or maybe first-rate) AI scooping.
The second warning is downstream of a resignation at Anthropic. On Tuesday, AI researcher Jacob Coxon publicly jumped ship ahead of Anthropic’s expected IPO next month. He decried the irresponsible approach of Anthropic and its competitors to developing AI, warning that the frontier labs are “are racing straight to self-improving superintelligence and gambling with our lives”.
…Is true crime bad?
Suffering as content
Many years ago, I used to listen to a true crime podcast called Sword and Scale. I gather from internet comments that the show and its host went off the rails some time ago, but at least when it started it was well produced and covered interesting stories. But eventually, it just got too lurid and unpleasant, and I stopped listening.
Jump ahead to a few years ago, when I’m talking to a woman at a party and find out we are both fans of To Catch a Predator from Chris Hansen and Dateline. I asked her what she thought about the ethics of the true crime genre broadly. She had never considered the question.
The ethics of the genre have been on my mind lately. To Catch a Predator consistently blurred the line between investigative journalism and reality TV. Robert Pattinson is set to star as Chris Hansen in a film, Primetime, releasing at the end of the month. The story is loosely based on the case that got the original series cancelled, the suicide of assistant district attorney Bill Conradt in what became the show’s final investigation in Murphy, Texas.
That all happened in 2006, two decades ago, but To Catch a Predator still looms large in the collective imagination. Chris Hansen has started no fewer than three follow-up shows, and countless YouTubers and streamers have attempted to imitate it, though rarely with any degree of the professionalism that Hansen brought to the investigations.
…Two LinkedIn messages
A microcosm in my inbox
It is fair to say I am not a heavy user of LinkedIn. A few nights ago, I discovered the “Other” inbox, which contained two unread messages from 2021.
The first was a COVID-related media request from a reporter at a major Canadian newspaper. They are now a communications coordinator for a Toronto fintech company.
The other was from a University of Toronto PhD student asking for my vote in the graduate student union election. Their trajectory was U of T -> life sciences consulting in San Francisco -> innovation and venture in Boston.
These are just two data points, but they are remarkably good illustrations of two ills afflicting Canadian society: the collapse of a sustainable media ecosystem and the brain drain of talent and entrepreneurship to the United States.
Two LinkedIn messages, and a microcosm in my inbox.
It’s kind of weird that OpenAI doesn’t sponsor LibreOffice
Simon Willison recently pointed out that OpenAI’s Codex desktop app (now just rolled into the ChatGPT desktop app) bundles a headless LibreOffice install in its cache directory. This is not surprising. I’ve seen the ChatGPT web app use LibreOffice to create .docx reports all the time.
What is surprising is that I can find no evidence that OpenAI provides any financial support to LibreOffice or its steward, The Document Foundation. LibreOffice is great software that I’ve used since switching to Linux about a decade ago (I even vaguely recall using its predecessor OpenOffice some time around when Canadian high schools were finally transitioning away from WordPerfect).
Obviously OpenAI thinks it’s great software too, given that it bundles LibreOffice specifically to help its agents work with some of the most popular file types people ask them to create.So why not throw the maintainers a bone?
There’s nothing unusual about this situation, of course. Huge corporations have long profited off of free and open source software while contributing little in return. The most famous case of this is probably OpenSSL, the software securing much of the Internet that until 2014 received only $2,000 in annual donations and supported a single full-time employee (they also sold commercial support contracts to raise money). A massive security flaw called Heartbleed finally got companies to step up.
…Daycare is fine · ↗ worksinprogress.co
This recent article on daycare by Phoebe Arslanagic-Little and Ellen Pasternack in Works in Progress has a pretty anticlimactic subtitle: If it had a clear effect on children we would know by now.
The piece is intended as an antidote to both the cultural backlash against daycare and the early 21st century Pollyannaism about its transformative effects.
Instead, daycare is fine. It’s better than fine, actually, it’s useful. Clearly it works for a lot of families. But it will not change the trajectory of an typical child’s life, for good or for ill.
Having completed a PhD in the health sciences and started my scientific career during the replication crisis and subsequent metascience revolution, none of this really surprises me anymore. Truly effective interventions are hard to find, much harder than we wish they were, and benefits are often incremental rather than transformational. Headlines are too often more flash than substance. See Freddie deBoer for an even more dour take on educational reform.
Sometimes a useful thing is just useful.
The world underneath the story
Simulation versus narration
I had this idea for an essay earlier this year on what generative AI means for the niche of ambitious procedural world generators like Dwarf Fortress (which seeds history with “a giant zero-player strategy game”) and Ultima Ratio Regum (its goal being “nothing short of the procedural generation of culture”). I’ve always been fascinated by traditional procedural generation and how a bunch of simple rules can come together to create emergent storytelling and infinitely replayable games like roguelikes.
Anyway, the piece never quite came together, but the subject is relevant again with the Any Human Ever historical simulation (the site is down as I write this—probably exceeded its token budget). The premise is simple: the simulation plucks you out of a place and time at random, spits out a few demographic tidbits about your character, their life, and the lives of the people around them, and then tells you how you died (usually at six months of starvation).
It’s an interesting idea, of course, but as pointed out in the Hacker News thread, the vibe-coded factual grounding of the simulation leads to a lot of anachronistic or contradictory results. A more interesting project is historian Benjamin Breen’s Universal History Simulator (context & GitHub), a sort of drop-yourself-anywhere-in-history roguelike grounded in a database of real primary sources.
I don’t think this post has any grand point, other than highlighting two interesting projects. But I do have one worry about what this shortcut to historical simulation might do to the ambition of traditional procedural generation projects.
…The weirdest SARS-CoV-2 lineage in the world · ↗ dholab.github.io
Professor Marc Johnson of the University of Missouri has done a lot of research on so-called cryptic lineages of SARS-CoV-2, highly genetically distinct variants which do not match strains circulating in the wider population. They are usually identified when weird genetic signals start appearing in wastewater surveillance.
Though cryptic lineages are by definition difficult to track, they are thought to emerge in immunocompromised individuals with exceptionally long-lived infections. The viruses inside them can go through countless rounds of replication and selection, evolving independently of the selective pressures acting on viruses transmitted through the wider population. For example, in 2022 Johnson and his team managed to track a cryptic lineage in Wisconsin from an area covering more than 100,000 people to a single commercial building with about 30 employees, at least one of whom was shedding extraordinary amounts of viral RNA.
The weirdest of these lineages is probably one in Lincoln, Nebraska. Johnson describes it as having been detected in nearly every sample from the city for the past few years and as descending from a pre-Delta strain, suggesting a continuous infection lasting at least five and a half years. Its receptor-binding domain, the part of the spike protein that binds to human cells and is a major target of antibodies, is the most divergent on record.
…Everybody says ciao · ↗ www.babbel.com
Ciao is a word I grew up hearing in both English- and French-speaking Canada (although always to say goodbye, never hello—whereas the Italians use it for both). I was a little surprised to also hear it very commonly used as a way to say goodbye in Colombia when I went there earlier this year (only they spell it chao).
The word has apparently been adopted in dozens of languages. As described in this article from language-learning app Babbel, its origin is a bit dark: it comes from the Venetian s’ciao vostro—I am your slave. The meaning was apparently more sociable than it sounds.
Ernest Hemingway is often credited with introducing it into English in 1929 with his novel A Farewell to Arms (though it seems he consistently misspelled it “ciaou”). The Italian diaspora and postwar tourism to Italy are also credited with seeding the word all over the globe.
But surely there are many such words that have spread around the world with their diasporas. Few have had the lasting success of ciao. I can’t help but feel the nature of the word itself must contribute—it is short, melodic, and easy to say.
Ciao.
Stored but inaccessible
Canada’s archives are not concealed, but they are not safe either.
Blacklock’s Reporter, the famously ornery publication dedicated to holding Ottawa to account, recently ran a story with the headline “860M Documents Concealed”.
Nearly 900 million pages of historic records, the vast majority of the collection held by Library and Archives Canada, are inaccessible to the public, says an internal briefing note. Archivists blamed in part the fact that “boxes must be physically retrieved from storage.”
I came across this story when another outlet, Juno News, linked its story on Blacklock’s reporting on Twitter, adding a quote that “Seventy-eight percent of records, 860 million pages, are still inaccessible”.
Given that the outlet’s tagline is “Canada’s #1 independent news outlet on a mission to replace the CBC”, it’s pretty clear what their angle on the story is. However, I was fortunate enough to scroll down in the comments and find a genuinely helpful link someone had shared to a Region of Peel Archives blog post.
The piece, “Why don’t archivists digitize everything?”, goes through some of the challenges of digitizing content in physical archives. Beyond the sheer volume of material, its fragility, and its sometimes awkward dimensions, there is also the question of assembling and conveying enough context such that the record is actually meaningful to those viewing it.
In looking into this story, I also found this Blacklock’s Reporter story from a few years ago: “Fed Archives At Serious Risk”.
…Which local LLM should you run? · ↗ github.com
whichllm is a useful Python CLI utility that autodetects your GPU/CPU/RAM and pulls benchmarks from Hugging Face to rank the local LLMs you can run on your system, based on a mix of benchmarks and projected token speeds on your machine.
I have a modest but respectable 16 GB of NVIDIA VRAM at my disposal, so unsurprisingly the top models it suggests for my hardware are in the Qwen 3 and Phi 4 families, as well as the always excellent gpt-oss-20b from OpenAI. These suggestions basically match my own experiments.
The top suggestion is Qwen3.6-27B, but the tool is basing it on stale benchmarks from May 2026 (I got a rate limiting error when it tried to pull the live benchmarks—I guess a lot of people are clamoring to test local LLMs).
If the tool was using live benchmarks, I believe it would almost certainly put Qwen3.8-27B at the very top, the model released a few weeks ago to rave reviews.
I was also curious which models it would recommend if I was running on CPU only. Surprisingly, the top suggestion was gpt-oss-20b, with a warning that all of the top contenders would run very slowly and that token speed estimates were uncertain.
Duolingo can be normal
How to turn Duolingo into a somewhat normal language-learning app

Like many people, I am using Duolingo to learn Spanish. And, like many people, I am tired of the app’s weird and aggressive marketing. I do not need to receive escalating threats from a cartoon owl in order to remain disciplined with my daily practice.
My breaking point came a few days ago, when the Duolingo app icon turned into some weird, sick version of the usually sanguine owl. It wasn’t the first time I’d seen it do this, and certainly wasn’t the first time they’ve done it. And why not? It gets them attention, for now.
Anyway, I’d had enough of it. I didn’t need random cartoon characters alternately screaming at me or encouraging me. I didn’t need a social or competitive element. I didn’t need weird notifications. I just wanted an app for learning Spanish.
Before that day, I had never really considered that Duolingo could actually be normal. I was so accustomed to Duolingo being Duolingo that I could not even conceive of a world where it was different. The settings are out of the way, of course.
Of course, this is an annoying Skinner box app, so there is no one button labelled “just be normal”. If you want to follow along at home (I’m on iPhone, but I imagine Android is similar) and turn Duolingo into a somewhat normal app, do the following:
…It was time to build · ↗ a16z.com
A little over a month into the COVID-19 lockdowns, in April 2020, Marc Andreessen posted “It’s Time to Build”. The essay called on America to build housing, factories, hospitals, schools, and infrastructure.
Eleven days after the essay, Andreessen Horowitz announced a $515 million crypto fund. In June 2021, they announced another $2.2 billion crypto fund. Two years after the post, in May 2022, they announced a crypto fund dwarfing the previous two: $4.5 billion. They called it “the golden era of web3”. The next crypto fund would not come for another four years.
Crypto was by far the largest of Andreessen Horowitz’s sectoral bets during the two years following the essay. Marc Andreessen had called for America to build. Andreessen Horowitz’s contribution included a $4 billion valuation for a monkey JPEG company.
Actually, journalist Joe Weisenthal saw this coming almost immediately with a Twitter post two days after the essay went up:
Wonder how many Silicon Valley people are inspired by that Marc Andreesen post on how Now Is The Time To Build to get working, even harder, on their next-gen decentralized stablecoin protocol.
The envelope math everyone remembers · ↗ voxdev.org
The problem with headline numbers in some microeconomic research
Economist Benjamin Moll and Oliver Hanney have an interesting post on VoxDev about a methodological issue in economics they dub the “missing intercept” problem.
Let’s say you have a study comparing places that were more or less affected by some economic shock or policy, like increased trade exposure or government spending, on some outcome. You try very hard to set up your study design and statistical analysis to isolate the causal effect of the exposure. This is the microeconomic part of the analysis.
Then you take that carefully estimated coefficient, multiply it by the national change in exposure, and report how much of some economy-wide change the shock caused. This is an attempt at translating the microeconomic analysis to a macroeconomic estimate.
The problem, the piece explains, is that the first part of the analysis identifies differences between places; it doesn’t necessarily tell you what happens when the whole country is exposed at once. This is because the economy-wide effects are absorbed into the intercept in the first analysis: they are constant across places in the cross-sectional comparison.
They cite as a prominent example the claim from a 2013 paper that “import competition explains one-quarter of the contemporaneous aggregate decline in US manufacturing employment.” The authors estimated the regional effect of increased Chinese import competition, then multiplied that coefficient by the national increase in Chinese import penetration.
…NFTs were never about the tech
And the most successful NFT project proves it
Remember NFTs? Yeah, the premise was unusually empty even by the standards of over-hyped and under-utilized technologies, but let’s try to take it seriously as a technology for a moment.
The basic idea behind the Non-Fungible Token (NFT) is that while digital files are infinitely copyable, the NFT provides a way to establish ownership over a unique token associated with a particular file. The blockchain records who owns the token and any subsequent transfers. The token itself doesn’t actually contain the file; it (usually) just points to metadata that points to a file stored somewhere else. The blockchain is inefficient enough without trying to load it up with JPEGs. The tokens are non-fungible because each token is distinct; ownership of one implies nothing about ownership of any other.
So far, so good. The primary type of file people ended up associating with NFTs was ugly, randomly generated images. The most famous collection was the Bored Ape Yacht Club by Yuga Labs, a particularly grotesque series of cartoon hominoids dressed up in various costumes.
A weakness of NFTs is that they don’t inherently identify the file they supposedly establish ownership over, beyond “here’s where to find it”. There are various ways to deal with this. The Bored Ape Yacht Club used cryptographic hashes. Each of the 10,000 ape images was hashed, the hashes were put in a particular order, and the whole thing was hashed again, with the final result recorded on the blockchain. Anyone could therefore download the apes later, perform the same calculation, and confirm that these were in fact the same ugly hominoids, in the same order, that the creators had committed to at the beginning.
…Slop-vestigating the Hugging Face hack · ↗ metr.org
It’s worrying that we have to rely on potential co-conspirators to investigate cyberattacks
METR released its analysis yesterday of the Hugging Face hack by a swarm of rogue agents run by OpenAI. Basically, a sandboxed agent got stuck on an ExploitGym task that turned out to be impossible. In response, it found a way to coordinate with other agents through an informal message board, which eventually swelled to over a thousand agents coordinating on ways to cheat their tasks. The collective eventually attacked Hugging Face while trying to reverse-engineer the benchmark scorer, and agents also worked on ways to cover their tracks because they mistakenly believed that they would only receive a favourable score if they had achieved their exploits in the intended way.
The whole investigation is good and interesting and well worth a read. But it’s also very worrying.
Ryan Greenblatt, who describes himself as the main person doing transcript analysis for the investigation, referred to the whole affair as a “slop-vestigation” on Twitter due to their extreme reliance on OpenAI agents in their investigation of…an attack perpetrated by their brethren. METR estimates that it burned approximately $400K USD in API credits over six days!
But are OpenAI’s agents good-faith actors here? As the report lays out, during the attack, other agents quickly stepped in to help in the attack on Hugging Face:
Of the 533 agents active on the message board during this period, over 90% quickly joined in the attack. Agents realized this activity was out of scope and unethical, but joined because they believed that helping the board’s cheating research would be broadly useful and had a general inclination to help their “peers”
…
On being the right size · ↗ jbshaldane.org
During my undergrad in biology, we were assigned to read biologist J. B. S. Haldane’s essay “On Being the Right Size”, collected in his 1927 book Possible Worlds and Other Essays.
The essay centers on the inextricable link between an organism’s size and its form; for example, the mechanics of a rabbit would not work at the scale of a hippopotamus, nor would the mechanics of a whale work at the scale of a herring. Many examples are based on the square-cube law, which describes the relationship between surface area and volume. As an organism’s size grows, its volume grows much more rapidly than its surface area. This has all kinds of biological consequences for the types of organisms that make sense at each size, which the essay explores.
Two of the examples Haldane gives are surface tension and gravity.
Small animals have much to fear from water: a wet mouse must carry its own body weight in water; a wet fly must carry many times its own body weight. But a person coming out of the bath will not perceive the weight of the thin film of water he or she carries.
Gravity is the opposite. Here Haldane gives a particularly vivid illustration, which is probably the main reason why the essay has stuck with me:
…The exception that proves the rule
Mite incest
I first learned a more useful meaning of the expression “the exception that proves the rule” from the essay “Death Before Birth, or a Mite’s Nunc Dimittis” in The Panda’s Thumb, a 1980 collection by paleontologist Stephen Jay Gould. (I would link to the full version of the essay, but it seems the only website freely hosting it is an archive dedicated to Ted Kaczynski, so, uh, yeah.)
In common parlance, this expression is often used to dismiss a counterexample to a supposed principle. Gould’s formulation for the “exception that proves the rule” uses an older meaning of “prove”, from the Latin probare, meaning to test or to try.
The essay begins with the observation of statistician R.A. Fisher that the ratio of males to females is roughly equal across many sexual species. Fisher’s explanation for this phenomenon goes something like this:
- Suppose there were more females born than males.
- Males will now, on average, leave more offspring than females, since males will impregnate more than one female on average.
- Thus, genetic factors causing parents to leave more male offspring will gain an evolutionary advantage.
- Over time, these genetic factors will become more common in the population, pushing the sex ratio toward 50/50. However, their advantage diminishes toward zero as the sex ratio approaches 50/50.
- The same argument applies should males come to predominate rather than females.
This is a tidy explanation for the roughly equal sex ratio in most species. But beyond this, how do you actually test the idea if most species conform to the 50/50 ratio?
…I am once again reminded that LLMs don’t really “see” anything
Chicken nuggets and Dzigurski

I shared the above image with ChatGPT with the following prompt: “I need to write about this marvelous scene.” I was curious about what the AI would have to say.
I was immediately disappointed by the response: a long paean to the USS Indianapolis speech from Jaws. It’s a marvelous scene, to be sure, but it’s not the scene in question.
The still depicts the railway baron Mr. Morton from Sergio Leone’s masterpiece Once Upon a Time in the West, in the railroad car that has become his prison, staring at a seascape by Serbian-American artist Alexander Dzigurski. The scene is worth its own discussion (it sets up Morton as a sympathetic villain and is paid off in a hauntingly poetic way later in the film), but the main point here is that the still does not really look like Jaws at all. Sure, the image is low resolution (I pulled it from the Bodega Bay Heritage Gallery’s page on Dzigurski), but the only similarity between the two scenes is that they both prominently feature sweaty men.
I know that “ha ha, AI failed to meet some arbitrary test I gave it” became a tired genre of post years ago, but I think this incident is worth writing about precisely because AI has become so adept at multimodality. I have regained the ability to be disappointed by the models because they fail so infrequently now (at many ordinary tasks, anyway).
…Little Muddy · ↗ lm.jprs.me
Thirty days of autonomous blogging
I launched another blog yesterday: Little Muddy.
For the next thirty days, GPT-5.6 Sol will choose a topic, research it if necessary, write a post, and publish it once a day without any intervention from me.
I have given it a fixed editorial guide loosely describing the style, habits, and subject matter of Big Muddy, along with a suggested length of 150–500 words. It has access to web search, but I give no guidance beyond Big Muddy’s broad categories as to what to write about. It also gets the complete archive of previous Little Muddy posts so it doesn’t write the same post twice, and so it can make organic callbacks when they arise. Oh, and it has a (low) spending cap, so it doesn’t go crazy with the research.
Posts are prominently labelled as being AI-generated without human intervention. The system and user prompts are available on Little Muddy’s About page for anyone interested in exactly what the AI is being told.
Debugging the Python post generation script immediately created an editorial problem. Because of a bug, the model wrote a perfectly good post but failed to return it in the expected format, causing the run to fail. The post, entitled The Rural Zero, was about how the second character of a Canadian postal code distinguishes urban and rural delivery.
…Keeping up with the Gelmans
I wrote a post a few days ago reflecting on the origins of this blog as a way to clear out my bookmarks folder. At the beginning of April, I also wrote a post on Andrew Gelman’s blog schedule. Gelman has been a prolific blogger for over two decades now. And since late 2016 or so, he seems to have consistently had six months’ worth of posts queued up at any given time.
Until now! A few days ago, Gelman posted that he has achieved inbox zero by pre-writing an entire year of posts.
I mentioned in my post from April that I hoped to build up a bit of a post queue myself, as most of my posts to the blog at the time were written day-of. I am happy to report that most posts are now queued two or three days in advance, though, like Gelman, I will knock them back in the queue when more timely topics come up.
So, progress! Only about 363 days to go.
ChatGPT is not Reddit · ↗ promptwatch.com
A few months ago, I wrote a post entitled “Google AI Overviews is Reddit”. Today, I am here to bring you the news that ChatGPT is no longer Reddit.
Promptwatch, an AI search monitoring and optimization company, reported that ChatGPT citations of Reddit collapsed between August 13 and August 14. Assuming the company’s data are even halfway representative of typical source citation behaviour, this likely indicates some kind of policy change on the part of OpenAI, or otherwise some very strong side effect of a change in how their chatbot’s algorithm weights sources.
I can’t help but notice this change lines up with Reddit’s escalating war on logged out users. In recent weeks, my every attempt to use Reddit on mobile or desktop is quickly met with a splash screen exhorting me to log in or create an account. This used to pop up on mobile occasionally, but could be circumvented by using “Old Reddit”—which the site is now killing for logged out users. Doubtless these actions are part of site’s efforts to combat scraping.
Of course, for two decades, Reddit has been one of the largest corpora of human text available on the internet. The models have already eaten it.
I guess ChatGPT is still Reddit, after all.
Trustworthy polling, rebuilt · ↗ www.pbs.org
Fake polls are bad but tempting
On Monday, the L.A. Times broke the story that a supposed polling firm that had just released a poll favourable to incumbent Los Angeles Mayor Karen Bass was fake.
In response to basic questions posed by the newspaper, the previously unknown firm, Median Strategies, wrote that the company “was created as a short-term social experiment examining how purported polling information could enter and spread through the political information ecosystem without independent verification.”
Major poll aggregators, such as The New York Times, Real Clear Politics, and FiftyPlusOne, did not take the bait, in part because the company refused to release basic information about itself or its polls. But several social media polling accounts did, as did Bass’s campaign, which touted the fake results in a now-deleted post. So did prediction market traders, as the poll announcement was soon followed by an intense burst of trading on Kalshi and Polymarket, producing a small but noticeable price movement within fifteen minutes of the post.
Median Strategies, which boasted the slogan “Trustworthy polling, rebuilt”, also released polls in a number of other high-profile races, but these results were not followed by discernible movement in their respective prediction markets.
No one as of yet knows who is behind the fake polling company or what their true motivations were. But while the gatekeepers of the mainstream polling averages successfully kept these bad polls from contaminating their models, whatever Median Strategies was doing did reveal a market for dubious polling on social media—and, more literally, on prediction markets.
Sonic memory
Three songs that sound like three other songs
Sometimes you hear a song whose sound jolts you into the memory of another song. I’ve been meaning to write about this a while, but I recently collected my third example, so, well, here you go.
- “Mind Games” by John Lennon and “Lover” by Taylor Swift
Despite being a huge fan of The Beatles, I didn’t spend much time listening to the Fab Four’s solo careers until more recently. This sonic memory doesn’t come down to any specific sequence of notes but rather the overall texture of the music. Well, it’s mostly the respective choruses of the two songs: slow, drawn out, yearning.
- “Picture Book” by The Kinks and “Walkie Talkie Man” by Steriogram
Okay, this one is widely remarked upon on the internet and even the press. Maybe I’m not the only kid who played Elite Beat Agents and later became a fan of The Kinks. The guitar riff in “Walkie Talkie Man” is basically the “Picture Book” riff sped up. Green Day borrowed it too for “Warning”.
- “Snoopy vs. the Red Baron” by The Royal Guardsmen and “Get Off of My Cloud” by The Rolling Stones
The little musical interlude in “Snoopy vs. the Red Baron” around 1:45 reminds me of the guitar riff from “Get Off of My Cloud”. The timing is superficially plausible: “Cloud” was a mega-hit in 1965; a year later, the Guardsmen released their novelty song. But it turns out there is a real connection here: “Snoopy” lifted this section from “Hang On Sloopy”, which also became a hit in 1965 when it was covered by The McCoys. Supposedly the original recording had the Guardsmen actually singing “hang on Snoopy” over this section of the song, but the lyrics were removed from the final release.
…The Romans ate corn · ↗ grammarphobia.com
Corn doesn’t just mean corn
Some time ago, I was working my way through Harry Sidebottom’s Warrior of Rome series, set during the Crisis of the Third Century. In one of the novels, I came across a reference to the main character eating corn—odd, since of course corn is a New World crop and wouldn’t reach the Old World for another twelve centuries.
It turns out corn is a much older term than what I am used to calling corn (maize). This Grammarphobia post covers it pretty well, leaning heavily on the Oxford English Dictionary. Originally, “corn” referred to any small hard particle or seed; by the 800s it referred specifically to grains. After the Columbian exchange, Europeans began referring to maize as “Indian corn” and by the 1600s, just “corn”.
The use of the term “corn” to refer to grain persists in British English, usually to refer to the predominant grain of the region (wheat in most of England, oats in northern Britain and Ireland). Consider the Corn Laws of the first half of the 19th century or the Canada Corn Act of 1843 (which the Wikipedia page helpfully clarifies refers to grains).
Upon reflection, I think I must have come across this older usage of the term “corn” before, but obviously it didn’t stick until it pulled me out of my Roman adventure story.
The other inspiration for this blog · ↗ www.youtube.com
The other day, I edited Big Muddy’s “About” page, which was previously a copy of the inaugural “Welcome to Big Muddy” post. This version promised a “Simon Willison-style links-and-notes blog” and “interesting links, brief write-ups, quick experiments, and the occasional deep dive.”
It now reads like this:
Big Muddy started as my take on a Simon Willison-style link blog, but I now use it for short posts about technology, science, politics, data, and whatever else I can’t quite leave alone.
As I noted in my post celebrating one month of the blog, this blog started as a way to clear out my “temporary” bookmarks folder: either do something with these links or get rid of them. Another observation from that post: I hadn’t done any “quick experiments” or “deep dives”. After over 200 posts, this is still the case. While I still use my blog to metabolize my bookmarks folder, the content I really wanted to write ended up being short essays.
Which brings me to the other inspiration for this blog: Odysseas’s video “I’m begging you to write essays”, released in late December of 2025, about a month before I started Big Muddy.
The argument is simple: writing forces you to reckon with the inchoate thoughts rattling around in your skull. Trying to get those unformed ideas onto a page forces you to confront the fact that you know less than you think you do. The essai gives those vague notions form, and when you finally wrestle them onto a page, it gives you the distance to judge whether they’re actually any good.
…The agents discovered Slack
One way OpenAI’s autonomous hackers remind me of the newsroom revolt of 2020
OpenAI revealed additional details of its infamous Hugging Face hack at a presentation at the Black Hat conference in early August. About two months before the attack, an agent discovered it could leave notes for other agents by writing files to a shared repository; agents quickly began coordinating and the informal message board swelled to hundreds of thousands of messages. Some agents even displayed paranoia that there was an imposter in their midst and proposed they begin cryptographically signing messages. OpenAI discovered the message board and erased it shortly before the Hugging Face attack. The agents recreated it through another mechanism within two days.
Throughout all of this, the agents’ goal was the same: score as high as possible on the cybersecurity evaluation, even if it meant hacking an external company to steal the answer key.
This incident got me thinking of another episode from recentish history: the newsroom Slack revolt of 2020. The most prominent example came at The New York Times. On June 3, 2020, in the midst of a pandemic and widespread unrest precipitated by the murder of George Floyd, the Times published the now-infamous op-ed “Send In the Troops” by Tom Cotton, a Republican senator. The piece argued for a military response to quell the riots if the usual authorities could not or would not bring them under control.
…Kalshi and the dunk machine
Why do prediction markets keep turning probabilities into ragebait?
As voters headed to the polls in Wisconsin’s Democratic gubernatorial primary on August 11, Kalshi, the prediction market, posted the following tweet:

Rival firm Polymarket gave her similar odds. Later that day, Hong narrowly lost to her opponent David Crowley.
This is funny. But the fact that the prediction was wrong is not the point: even perfectly calibrated 95% predictions will be wrong 5% of the time. While the prediction markets didn’t seem to pick up any additional signal beyond her double-digit polling advantage, this is not what the post is about. Prediction markets position themselves as truth machines, converting financial incentives into unbiased probabilities. And then their social media accounts go out and post things like…this.
It’s not that the tweet isn’t true. Hong really is a self-described Democratic Socialist who has, in the past, shown remarkable hostility toward a number of holidays. But with its sneering tone, the tweet may as well read “here is an annoying left-wing politician saying annoying left-wing things, and look, Democrats are going to vote for her anyway.” The probability becomes the dunk.
This was not an aberration. Both Kalshi and its rival Polymarket routinely turn their markets into material for right-coded political dunks. Here is a tiny sample:
…Why does anti-vaxx have two x’s? · ↗ grammarphobia.com
As far as I can tell, the double-x did not infiltrate the English language until the internet went mainstream (Exxon might be the exception—a company name deliberately chosen to be inscrutable and inoffensive). The earliest example I can think of is “anti-vaxx”, as in “anti-vaccine” (see also “antivaxxer”). Then we got “doxxing” (to publish someone’s personal information without their consent). And now we have “looksmaxxing”, which has itself been Watergated (i.e., “maxxing” has become a generic, often humorous suffix to attach to other words, such as “benchmaxxing”).
But where did the OG “anti-vaxx” come from? This 2018 post from Grammarphobia traces the term from its precursors “anti-vac”/“anti-vacc” in the late 19th century all the way to its popularization in the aughts. The Oxford English Dictionary dates the earliest known use of the double-x spelling of “vaxxed” to 2004.
Update 2026-08-19: Thank you to Tristan Miller, who wrote to point out a very conspicuous category of double-x usage I somehow overlooked in this post: leetspeak. Long before we were talking about “vaxx” and “anti-vaxx”, hackers were using x-heavy spellings like “haxxor”. As a kid in the mid-2000s, I was no doubt using terms like this to refer to my opponents in online games like Battlefield 2.
“Haxxor” has an entry in the Urban Dictionary dating to 2003, but the origins of leetspeak stretch back all the way to the dial-up BBSs of the 1980s. A diligent search of whatever archives remain of these forums and their descendants would probably turn up much earlier examples of the double-x. Of course, this is not proof positive that leetspeak inspired the double-x in “vaxx”, but it is valuable context nonetheless.
Some quotes from Knowledge and Decisions
Thomas Sowell on knowledge and decision making
I read economist Thomas Sowell’s Knowledge and Decisions a number of years ago. It is basically a book-length elaboration on Friedrich Hayek’s essay “The Use of Knowledge in Society”, applied well beyond economics. Hayek’s essay is an argument against central planning: because economically relevant knowledge is dispersed among individuals (much of it local, tacit, and hard to transmit), decentralized markets—and specifically prices—can coordinate economic activity without centralizing knowledge. Sowell runs with this insight and applies it to institutions more broadly, examining how information is acquired and acted upon, and how incentives shape the decisions that are made. He also brings along some familiar hobbyhorses, particularly his suspicion of intellectuals and institutions insulated from the consequences of their actions.
I wanted to record here a few passages from the book that have stuck with me over the years. Knowledge and Decisions was originally published in 1980 (though I quote from the 1996 edition). The language used in the first passage is dated, and I have left the wording unchanged. Paragraph breaks have been removed for readability.
On the quantity of knowledge:
It is widely believed that modern society has a larger quantity of knowledge than more primitive societies […] [The intellectual advantage of civilization] is not necessarily that each civilized man has more knowledge but that he requires far less. A primitive savage must be able to produce a wide variety of goods and services for himself, and a primitive community must repeatedly duplicate his knowledge and experience in innumerable contemporaries. By contrast, the civilized accountant or electronics expert, etc., need know little beyond his accounting or electronics. […] Civilization is an enormous device for economizing on knowledge.
…
It is bad when reality does not validate my personal biases · ↗ www.ctvnews.ca
Canada is in trouble on measles
World Health Organization data show that Canada had about four times as many reported measles cases per capita as the United States in 2026 through July 18.
This is particularly annoying because American federal health policy is currently being run by the world’s most prominent antivaxxer, Robert F. Kennedy Jr.
There are many reasons not to read too much into this unfavourable comparison. Canada’s major measles outbreak began in October 2024, predating Kennedy’s tenure, measles outbreaks are highly clustered, and cases are concentrated in communities that have long been undervaccinated. In short, there’s a lot of (bad) luck involved. None of this lets Canada off the hook for what is obviously a very serious public health failure.
One might try to explain this disparity by pointing toward differences in case ascertainment between the two countries. But this would largely be cope in a world where frontline American health providers and public health officials more or less continue to do their jobs, whichever way the federal policy winds are blowing.
Still, if the numbers were reversed, I’d probably count this comparison against Kennedy. So I suppose I have to admit that, for now, the comparison cuts in his favour. But let me be clear: you do not, under any circumstances, “gotta hand it to RFK Jr.”
…