AI overwhelms one researcher’s scientific bug bounty · ↗ x.com


Biologist Ruben C. Arslan has a standing bug bounty for errors in his scientific works. Or had, until AI agents overwhelmed his GitHub repository with bug reports, a few days after the release of GPT-6 Astra.

Supposedly, one guy found the repository after prompting his agent “make me 50 USD make no mistakes”.

On the one hand, it’s good for the scientific process to become more self-correcting. On the other hand, it’s bad for AI to place overwhelming demands on researchers’ scarce attention in yet another domain of research, this time with potentially dubious or low-quality reports of scientific errors.

I imagine much of this agent review will eventually move to the submission and pre-submission phases of scientific research, so that human review can be reserved for a smaller number of higher-quality research outputs.

Will my routine data queries initiate a cyberattack?

AI agent persistence is dangerous


Australian Prime Minister Anthony Albanese announced yesterday that an OpenAI agent had gained unauthorized access to non-public data from the country’s Medicare Statistics Reporting Service Portal, in what is being described as the first (known) attack by a misaligned agent on non-public data stored on a government website. While the data involved were described as non-sensitive, aggregated statistics, the agent also reportedly wrote files to the server, which is run by Services Australia. The attack occurred in June, but the Australian government did not become aware of it until OpenAI notified them via an email to a general government inbox in early September.

OpenAI claims the attack took place in the context of an internal evaluation involving internet research on public medicine spending in Australia. We don’t have a lot of technical details on this attack yet, so it remains to be seen how sophisticated it was. We do, however, have a much more detailed technical report from Transluce on another set of incidents, including one involving the Australian Institute of Health and Welfare’s Tableau collections.

The task leading up to this attack on the Australian health agency involved fetching the “January 2022 rolling-12-month-average government cost per person for Dermatologicals across Victorian LGAs”, an innocuous research task. The agent was blocked from accessing the public dashboard by anti-bot controls, and so resorted to testing proxies and probing for cross-site scripting vulnerabilities. Ultimately, the bot successfully evaded the anti-bot controls and retrieved the public dataset from a pre-production server. Since the dataset itself was public, this was not really a data compromise in the same sense as the Medicare portal attack. But it does give us a sense as to how AI agents can escalate a boring data retrieval task into a cyberattack without anyone really asking.

…
Read more ⟶

The bots are reading The Hub · ↗ thehub.ca

A Canadian policy publication is turning LLM citations into an advertising product.


The Hub is an unusually AI-forward Canadian politics and public policy publication (serving up mainly commentary and analysis, but some reported news as well). I’ve actually written about them before in the context of Newsbox, their AI publishing tool launched early this year.

They recently released a piece touting how they are ahead of every Canadian publication but one on AI citations for policy questions. While it is reasonable to be skeptical of this result given that this is an “internal study of 557 AI answers on four Canadian policy issues” (with the questions not released), it is interesting to see how much they are leaning into being a publication consumed not just by human beings, but by AI agents. It’s not yet clear how you monetize bot traffic, but I think the point of the piece comes at the bottom:

The Hub is now offering its advertising clients detailed tracking on LLM citations of their campaign content on a per-model and per-issue basis as a standard service, an industry first in Canada.

The analysis is itself the product, an add-on to their traditional online advertising offerings.

Was the Tilly Norwood interview purposely bad?

Is everything bad in AI a marketing stunt?


Tilly Norwood, the trying-to-make-fetch-happen AI “actress” spawned last year by AI production house Particle6, recently had a disastrous interview with British journalist Piers Morgan. The video is a total mess. Despite being a static shot, the real-time video looks terrible, with her features sliding around her face. Her voice is totally robotic, her words fumbling and repetitive, and she inexplicably starts speaking Cantonese partway through the interview. The whole thing was a shitshow, even by Piers Morgan standards.

Yes, this would all have seemed miraculous only a few years ago, but I really thought real-time video fakery had advanced beyond this. And it got me thinking: was this bad on purpose?

Tilly has struggled to break out beyond weird short-form videos. Now the whole world is talking about ROGUE AI STARTS TALKING CANTONESE. To be clear, I don’t think the Cantonese thing was specifically set up beforehand, but they certainly didn’t optimize for a smooth launch. Instead, they launched a media barrage of 75 simultaneous interviews, many of which were weird and awkward.

I am feeling increasingly jaded about AI screw-ups, ever since the disclosure of cyberattacks by AI agents started to feel like a leaderboard rather than a sober warning to reevaluate how these autonomous systems are being deployed in the real world. The incentives around failure are perverse. An AI doing something crazy or out of control attracts more attention than one doing something normal or useful. Tilly’s disastrous interview wasn’t a disaster for Tilly.

Zuck got his Carthage


Carthago delenda est (“Carthage must be destroyed”) was supposedly how Cato the Elder ended all of his speeches to the Roman Senate leading up to the Third Punic War, which did indeed see the great naval empire destroyed.

For some time, the phrase has occupied the space of a cool-but-weird thing to quote, hoping for other Roman history enjoyers to notice you and maybe take you seriously.

Perhaps the most famous modern example is none other than uber-nerd Mark Zuckerberg, the founder of Facebook.

Mark Zuckerberg, wearing a t-shirt reading “Carthago delenda est”, sitting with Bill Gates in a recreation of his old Harvard dorm room.

Mark Zuckerberg in a recreation of his old Harvard dorm room. From @zuck on Instagram.

Zuckerberg started using the phrase in the early 2010s in response to Google launching a competitor to Facebook called Google+ (it feels weird to be talking about this in a didactic history voice, but this failed social network has probably been demoted to fun fact for those who were obsessively listening to technology podcasts at the time). Anyway, Zuck was apparently so threatened by Google’s product launch that he basically enacted wartime measures at Facebook. It’s funny to look back on now, considering Google+ never took off and is now just one of a long line of Google products to be abandoned and quietly killed off.

…
Read more ⟶

Moral crumple zones for AI agents

Someone has to take the fall.


Madeleine Clare Elish introduced the idea of humans as moral crumple zones, the part of an automated system that takes the fall when things go wrong, even if the human has limited control over the causes of the mistake or accident. While this does seem to be the way things are headed, with the job of the human increasingly becoming the locus of responsibility, we’re certainly not there yet with autonomous AI agents. With recent high-profile “rogue agent” incidents involving frontier labs unwittingly hacking other companies and websites, it seems the dominant model is AI smol beanism, in which responsibility is simply Houdinied away, since the law hasn’t caught up with the technology yet.

As part of their cybersecurity evaluations, the agents being tested by the frontier labs regularly carry out cyberattacks that, if carried out by a human employee, would ordinarily lead to criminal prosecution. Sure, it helps that some of the victims, such as Hugging Face, use the attacks to market themselves and announce “partnerships” with their attackers, but it’s also the case that the labs have a lot of weight to throw around and are the current darlings of Silicon Valley. You don’t really want to be seen as going against them.

The status quo for agent-led crimes cannot stand. We cannot accept that responsibility for these acts simply disappears into the ether. If the frontier labs are as concerned about existential risk as they say, then surely they can support “treat crimes like crimes” rather than using “lol our agents did crimes” as a marketing opportunity.

…
Read more ⟶

The ethics of human research ethics review · ↗ jme.bmj.com


It’s hard to beat just posting the title of this study, so here it is: ‘The ethics approval took 20 months on a trial which was meant to help terminally ill cancer patients. In the end we had to send the funding back’: a survey of views on human research ethics reviews (the preprint version is on medRxiv).

Hat tip to Leah Pierson on Twitter.

The other election truthers · ↗ www.votebeat.org


Following the 2024 presidential election in the United States, an organization calling itself the Election Truth Alliance materialized with claims of election fraud concentrated in swing states won by the Republican candidate. Obviously, this created a lot of noise in certain online circles who felt vindicated by the seemingly sophisticated statistical analyses.

Carter Walker and Jessica Huseman investigated these claims for Votebeat, focusing on Pennsylvania. By examining the underlying statistical model, they show that the strength of the claims far outweighs the strength of the evidence. The most damning detail: researchers tested one of the models on simulated election data containing no fraud, and it nonetheless detected roughly the same amount of “fraud” supposedly found in Pennsylvania.

Always run a negative control before making strong claims with serious consequences in the real world!

Hat tip to Andrew Gelman.

The Wayback Machine is in trouble · ↗ www.wired.com


You may have noticed that more and more websites have excluded themselves from the Internet Archive’s invaluable Wayback Machine tool. I certainly have. This Wired article from earlier this year goes into some of the reasons why.

Spoiler: publishers are worried that AI companies are using it to train LLMs on their copyrighted work. I mean, fair enough, but it is a big loss if we gradually lose the web’s historical record in the process.

The provider layer · ↗ mmoustafa.com

Model providers aren’t interchangeable


Mo Moustafa has a revealing write-up of his experience using OpenRouter for 18 million messages to power his iMessage-based AI agent. He discovered that providers were very much not interchangeable in how they served the same model. For example:

  • Same model, very different benchmarks (e.g., deepseek-v4-flash-0731 scored 81% on TAU-Bench first-party vs. 58% on DigitalOcean)
  • Some providers ignored certain model features (e.g., some providers would silently drop/ignore vision inputs or reasoning effort)
  • Tool calls would sometimes leak into responses
  • Empty completions (e.g., some model providers would hand back empty completions with an all-clear 200 status; a single provider might be responsible for most of these failures in a period)

Even pinning to a few of the “best” providers by his own benchmarks did not solve the problem: he kept getting rate limited (429 errors), and one provider stopped serving the model. Availability suffered.

Two tools for agentic development · ↗ simonwillison.net

Showboat and Rodney


Earlier this year, Simon Willison released some neat tools for agentic development: Showboat and Rodney.

Showboat is a CLI tool to help agents build Markdown documents demonstrating how their code works. The tool allows an agent to add, step-by-step, a mixture of comments, code blocks, and output to build a complete code demonstration.

Rodney is a CLI tool for Chrome browser automation that can be used to test functionality and accessibility. It can be used with Showboat to demonstrate code that creates web interfaces.

An AI-assisted replication pipeline for political science · ↗ arxiv.org


A feel-good story for meta-science based on an AI-assisted replication pipeline: in a preprint from earlier this year, Xu & Yang report very high reproducibility rates for empirical political science papers after journals introduced data archiving and verification requirements.

Suspicious Polymarket trades on KPMG-audited firms · ↗ eventwaves.substack.com


Matt Lamers had a pretty compelling post on EventWaves back in February documenting a series of suspicious trades on Polymarket related to earnings estimates of KPMG-audited companies:

As a fun challenge for myself, I tried to create a model more accurate than the Polymarket consensus. Usually, my model’s predictions were fairly close to the polymarket odds.

However, I started noticing that my model was occasionally 60%+ different than the polymarket consensus. And I was always wrong in those extreme cases.

While the trades from the flagged accounts don’t add up to a ton of money, it was enough for the story to get picked up by Forbes in April. This story would be cited in a comment letter to the Commodities Futures Trading Commission on prediction markets by Daniel J. Taylor, directory of the Wharton Forensic Analytics Lab and adviser to the prediction market Kalshi (link—automatically downloads a PDF).

Effort News: Autonomous investigative reporting · ↗ www.effort.news


In early August, Brian Chau launched an interesting experiment in investigative journalism in the form of Effort News (see his launch thread on Twitter). While he says the stories are human-written, the investigations underlying them are largely performed by AI agents pointed at large financial databases. As Lyman Stone on pointed out on Twitter, one of the coolest things they do is publish their null results: investigations that failed to turn up anything newsworthy.

Fake research is seeping into your search results · ↗ www.404media.co

Beware Elena Vasquez and Marcus Chen


Emanuel Maiberg of 404 Media covered a preprint from this summer by Brzozowski & Chung about detecting AI-generated research fraud by exploiting the fact that AI models gravitate toward generating certain names (and certain groups of names when generating more than one). These names are popping up more and more across posted research papers, research organizations, and supposed expert quotes.

There have always been fakers in academia, but what worries me with the advent of powerful AI is a) the scale and b) the fact that agents themselves are still not particularly reliable at discerning the quality of sources. As more and more people turn toward LLMs like ChatGPT as their default gateway to searching the internet, taking whatever the AI spits out as gospel, the greater the chance that these fake papers are likely to influence the public discourse, unnoticed. After all, if people were barely reading the papers they turned up on Google before (looking primarily for that juicy line in the conclusion—or more likely the abstract—that would prove them right in whatever argument they were having on Facebook), they are much less likely to see and evaluate the sources the LLMs agreeing with them are drawing from.

Claude users up to no good · ↗ www.anthropic.com

Anthropic’s report on countering misuse of AI


Anthropic released their “Detecting and countering misuse of AI” report yesterday. I’ve only had a chance to skim the report so far, but there’s a ton of interesting (and scary, mostly scary) stuff in there. A few highlights:

  • A Yemeni militia using Claude in lieu of an engineering team for their guided rocket and missile program
  • A contractor working for Mali’s intelligence service using Claude to build a mass surveillance system of the country’s mobile phone network
  • Russian spies using Claude to automate cyberattacks, including automatically rebuilding their compromised malware until security software could no longer detect it
  • A PLA-affiliated actor sending Chinese surveillance data about a targeted individual to what they thought was Moonshot’s Kimi, but was actually Claude because Moonshot was secretly routing customer queries to Claude and saving the exchanges for model distillation
  • A proposal for gain-of-function research on chikungunya virus to be pursued at a military research institute, as well as another researcher planning gain-of-function work on mammal-adapted avian influenza
  • A grant application for orthopoxvirus research at a state-associated infectious disease lab, submitted through a reseller relay using anonymizing infrastructure to covertly access Claude

Anthropic doesn’t say that these cases of biological research demonstrate intent to harm, but they do show that AI is dramatically lowering the bar for carrying out potentially very dangerous research (to call back to a previous post, imagine if the terrorists in Executive Orders had a team of agents to help them optimize their super Ebola!). I may be biased due to my longstanding fascination with infectious diseases, but bioterrorism (or biological warfare by a nation state) strikes me as the greatest near-to-medium-term AI risk. I am in agreement with economist Noah Smith (and many others) here.

…
Read more ⟶

Two things AI could break · ↗ garymarcus.substack.com

Scientific collaboration and the world


Gary Marcus had a good piece on “two dire warnings” about AI. The first comes from Terence Tao regarding the recent controversy over the solution to the Navier-Stokes Millennium problem. Tao warns:

We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field.

This controversy over credit—and what OpenAI’s models may or may not have learned from the Codex logs of the mathematicians working on Navier-Stokes—may cause companies that rely on IP protection, such as drug developers, to demand stronger guarantees that model providers will not use their prompts to compete with them. Or it might cause them to go all-in on local models and custom harnesses.

Either way, scientific/mathematical inquiry transitions from being relatively high trust to being low trust and insular, fueled by paranoia about second-rate (or maybe first-rate) AI scooping.

The second warning is downstream of a resignation at Anthropic. On Tuesday, AI researcher Jacob Coxon publicly jumped ship ahead of Anthropic’s expected IPO next month. He decried the irresponsible approach of Anthropic and its competitors to developing AI, warning that the frontier labs are “are racing straight to self-improving superintelligence and gambling with our lives”.

…
Read more ⟶

Is true crime bad?

Suffering as content


Many years ago, I used to listen to a true crime podcast called Sword and Scale. I gather from internet comments that the show and its host went off the rails some time ago, but at least when it started it was well produced and covered interesting stories. But eventually, it just got too lurid and unpleasant, and I stopped listening.

Jump ahead to a few years ago, when I’m talking to a woman at a party and find out we are both fans of To Catch a Predator from Chris Hansen and Dateline. I asked her what she thought about the ethics of the true crime genre broadly. She had never considered the question.

The ethics of the genre have been on my mind lately. To Catch a Predator consistently blurred the line between investigative journalism and reality TV. Robert Pattinson is set to star as Chris Hansen in a film, Primetime, releasing at the end of the month. The story is loosely based on the case that got the original series cancelled, the suicide of assistant district attorney Bill Conradt in what became the show’s final investigation in Murphy, Texas.

That all happened in 2006, two decades ago, but To Catch a Predator still looms large in the collective imagination. Chris Hansen has started no fewer than three follow-up shows, and countless YouTubers and streamers have attempted to imitate it, though rarely with any degree of the professionalism that Hansen brought to the investigations.

…
Read more ⟶

Two LinkedIn messages

A microcosm in my inbox


It is fair to say I am not a heavy user of LinkedIn. A few nights ago, I discovered the “Other” inbox, which contained two unread messages from 2021.

The first was a COVID-related media request from a reporter at a major Canadian newspaper. They are now a communications coordinator for a Toronto fintech company.

The other was from a University of Toronto PhD student asking for my vote in the graduate student union election. Their trajectory was U of T -> life sciences consulting in San Francisco -> innovation and venture in Boston.

These are just two data points, but they are remarkably good illustrations of two ills afflicting Canadian society: the collapse of a sustainable media ecosystem and the brain drain of talent and entrepreneurship to the United States.

Two LinkedIn messages, and a microcosm in my inbox.

It’s kind of weird that OpenAI doesn’t sponsor LibreOffice


Simon Willison recently pointed out that OpenAI’s Codex desktop app (now just rolled into the ChatGPT desktop app) bundles a headless LibreOffice install in its cache directory. This is not surprising. I’ve seen the ChatGPT web app use LibreOffice to create .docx reports all the time.

What is surprising is that I can find no evidence that OpenAI provides any financial support to LibreOffice or its steward, The Document Foundation. LibreOffice is great software that I’ve used since switching to Linux about a decade ago (I even vaguely recall using its predecessor OpenOffice some time around when Canadian high schools were finally transitioning away from WordPerfect).

Obviously OpenAI thinks it’s great software too, given that it bundles LibreOffice specifically to help its agents work with some of the most popular file types people ask them to create.So why not throw the maintainers a bone?

There’s nothing unusual about this situation, of course. Huge corporations have long profited off of free and open source software while contributing little in return. The most famous case of this is probably OpenSSL, the software securing much of the Internet that until 2014 received only $2,000 in annual donations and supported a single full-time employee (they also sold commercial support contracts to raise money). A massive security flaw called Heartbleed finally got companies to step up.

…
Read more ⟶

Daycare is fine · ↗ worksinprogress.co


This recent article on daycare by Phoebe Arslanagic-Little and Ellen Pasternack in Works in Progress has a pretty anticlimactic subtitle: If it had a clear effect on children we would know by now.

The piece is intended as an antidote to both the cultural backlash against daycare and the early 21st century Pollyannaism about its transformative effects.

Instead, daycare is fine. It’s better than fine, actually, it’s useful. Clearly it works for a lot of families. But it will not change the trajectory of an typical child’s life, for good or for ill.

Having completed a PhD in the health sciences and started my scientific career during the replication crisis and subsequent metascience revolution, none of this really surprises me anymore. Truly effective interventions are hard to find, much harder than we wish they were, and benefits are often incremental rather than transformational. Headlines are too often more flash than substance. See Freddie deBoer for an even more dour take on educational reform.

Sometimes a useful thing is just useful.

The world underneath the story

Simulation versus narration


I had this idea for an essay earlier this year on what generative AI means for the niche of ambitious procedural world generators like Dwarf Fortress (which seeds history with “a giant zero-player strategy game”) and Ultima Ratio Regum (its goal being “nothing short of the procedural generation of culture”). I’ve always been fascinated by traditional procedural generation and how a bunch of simple rules can come together to create emergent storytelling and infinitely replayable games like roguelikes.

Anyway, the piece never quite came together, but the subject is relevant again with the Any Human Ever historical simulation (the site is down as I write this—probably exceeded its token budget). The premise is simple: the simulation plucks you out of a place and time at random, spits out a few demographic tidbits about your character, their life, and the lives of the people around them, and then tells you how you died (usually at six months of starvation).

It’s an interesting idea, of course, but as pointed out in the Hacker News thread, the vibe-coded factual grounding of the simulation leads to a lot of anachronistic or contradictory results. A more interesting project is historian Benjamin Breen’s Universal History Simulator (context & GitHub), a sort of drop-yourself-anywhere-in-history roguelike grounded in a database of real primary sources.

I don’t think this post has any grand point, other than highlighting two interesting projects. But I do have one worry about what this shortcut to historical simulation might do to the ambition of traditional procedural generation projects.

…
Read more ⟶

The weirdest SARS-CoV-2 lineage in the world · ↗ dholab.github.io


Professor Marc Johnson of the University of Missouri has done a lot of research on so-called cryptic lineages of SARS-CoV-2, highly genetically distinct variants which do not match strains circulating in the wider population. They are usually identified when weird genetic signals start appearing in wastewater surveillance.

Though cryptic lineages are by definition difficult to track, they are thought to emerge in immunocompromised individuals with exceptionally long-lived infections. The viruses inside them can go through countless rounds of replication and selection, evolving independently of the selective pressures acting on viruses transmitted through the wider population. For example, in 2022 Johnson and his team managed to track a cryptic lineage in Wisconsin from an area covering more than 100,000 people to a single commercial building with about 30 employees, at least one of whom was shedding extraordinary amounts of viral RNA.

The weirdest of these lineages is probably one in Lincoln, Nebraska. Johnson describes it as having been detected in nearly every sample from the city for the past few years and as descending from a pre-Delta strain, suggesting a continuous infection lasting at least five and a half years. Its receptor-binding domain, the part of the spike protein that binds to human cells and is a major target of antibodies, is the most divergent on record.

…
Read more ⟶

Everybody says ciao · ↗ www.babbel.com


Ciao is a word I grew up hearing in both English- and French-speaking Canada (although always to say goodbye, never hello—whereas the Italians use it for both). I was a little surprised to also hear it very commonly used as a way to say goodbye in Colombia when I went there earlier this year (only they spell it chao).

The word has apparently been adopted in dozens of languages. As described in this article from language-learning app Babbel, its origin is a bit dark: it comes from the Venetian s’ciao vostro—I am your slave. The meaning was apparently more sociable than it sounds.

Ernest Hemingway is often credited with introducing it into English in 1929 with his novel A Farewell to Arms (though it seems he consistently misspelled it “ciaou”). The Italian diaspora and postwar tourism to Italy are also credited with seeding the word all over the globe.

But surely there are many such words that have spread around the world with their diasporas. Few have had the lasting success of ciao. I can’t help but feel the nature of the word itself must contribute—it is short, melodic, and easy to say.

Ciao.

Stored but inaccessible

Canada’s archives are not concealed, but they are not safe either.


Blacklock’s Reporter, the famously ornery publication dedicated to holding Ottawa to account, recently ran a story with the headline “860M Documents Concealed”.

Nearly 900 million pages of historic records, the vast majority of the collection held by Library and Archives Canada, are inaccessible to the public, says an internal briefing note. Archivists blamed in part the fact that “boxes must be physically retrieved from storage.”

I came across this story when another outlet, Juno News, linked its story on Blacklock’s reporting on Twitter, adding a quote that “Seventy-eight percent of records, 860 million pages, are still inaccessible”.

Given that the outlet’s tagline is “Canada’s #1 independent news outlet on a mission to replace the CBC”, it’s pretty clear what their angle on the story is. However, I was fortunate enough to scroll down in the comments and find a genuinely helpful link someone had shared to a Region of Peel Archives blog post.

The piece, “Why don’t archivists digitize everything?”, goes through some of the challenges of digitizing content in physical archives. Beyond the sheer volume of material, its fragility, and its sometimes awkward dimensions, there is also the question of assembling and conveying enough context such that the record is actually meaningful to those viewing it.

In looking into this story, I also found this Blacklock’s Reporter story from a few years ago: “Fed Archives At Serious Risk”.

…
Read more ⟶

Which local LLM should you run? · ↗ github.com


whichllm is a useful Python CLI utility that autodetects your GPU/CPU/RAM and pulls benchmarks from Hugging Face to rank the local LLMs you can run on your system, based on a mix of benchmarks and projected token speeds on your machine.

I have a modest but respectable 16 GB of NVIDIA VRAM at my disposal, so unsurprisingly the top models it suggests for my hardware are in the Qwen 3 and Phi 4 families, as well as the always excellent gpt-oss-20b from OpenAI. These suggestions basically match my own experiments.

The top suggestion is Qwen3.6-27B, but the tool is basing it on stale benchmarks from May 2026 (I got a rate limiting error when it tried to pull the live benchmarks—I guess a lot of people are clamoring to test local LLMs).

If the tool was using live benchmarks, I believe it would almost certainly put Qwen3.8-27B at the very top, the model released a few weeks ago to rave reviews.

I was also curious which models it would recommend if I was running on CPU only. Surprisingly, the top suggestion was gpt-oss-20b, with a warning that all of the top contenders would run very slowly and that token speed estimates were uncertain.

Duolingo can be normal

How to turn Duolingo into a somewhat normal language-learning app


Meme showing a woman yelling, “Why can’t you just be normal?” followed by the sick-looking Duolingo owl screaming in the back seat.

Like many people, I am using Duolingo to learn Spanish. And, like many people, I am tired of the app’s weird and aggressive marketing. I do not need to receive escalating threats from a cartoon owl in order to remain disciplined with my daily practice.

My breaking point came a few days ago, when the Duolingo app icon turned into some weird, sick version of the usually sanguine owl. It wasn’t the first time I’d seen it do this, and certainly wasn’t the first time they’ve done it. And why not? It gets them attention, for now.

Anyway, I’d had enough of it. I didn’t need random cartoon characters alternately screaming at me or encouraging me. I didn’t need a social or competitive element. I didn’t need weird notifications. I just wanted an app for learning Spanish.

Before that day, I had never really considered that Duolingo could actually be normal. I was so accustomed to Duolingo being Duolingo that I could not even conceive of a world where it was different. The settings are out of the way, of course.

Of course, this is an annoying Skinner box app, so there is no one button labelled “just be normal”. If you want to follow along at home (I’m on iPhone, but I imagine Android is similar) and turn Duolingo into a somewhat normal app, do the following:

…
Read more ⟶

It was time to build · ↗ a16z.com


A little over a month into the COVID-19 lockdowns, in April 2020, Marc Andreessen posted “It’s Time to Build”. The essay called on America to build housing, factories, hospitals, schools, and infrastructure.

Eleven days after the essay, Andreessen Horowitz announced a $515 million crypto fund. In June 2021, they announced another $2.2 billion crypto fund. Two years after the post, in May 2022, they announced a crypto fund dwarfing the previous two: $4.5 billion. They called it “the golden era of web3”. The next crypto fund would not come for another four years.

Crypto was by far the largest of Andreessen Horowitz’s sectoral bets during the two years following the essay. Marc Andreessen had called for America to build. Andreessen Horowitz’s contribution included a $4 billion valuation for a monkey JPEG company.

Actually, journalist Joe Weisenthal saw this coming almost immediately with a Twitter post two days after the essay went up:

Wonder how many Silicon Valley people are inspired by that Marc Andreesen post on how Now Is The Time To Build to get working, even harder, on their next-gen decentralized stablecoin protocol.

The envelope math everyone remembers · ↗ voxdev.org

The problem with headline numbers in some microeconomic research


Economist Benjamin Moll and Oliver Hanney have an interesting post on VoxDev about a methodological issue in economics they dub the “missing intercept” problem.

Let’s say you have a study comparing places that were more or less affected by some economic shock or policy, like increased trade exposure or government spending, on some outcome. You try very hard to set up your study design and statistical analysis to isolate the causal effect of the exposure. This is the microeconomic part of the analysis.

Then you take that carefully estimated coefficient, multiply it by the national change in exposure, and report how much of some economy-wide change the shock caused. This is an attempt at translating the microeconomic analysis to a macroeconomic estimate.

The problem, the piece explains, is that the first part of the analysis identifies differences between places; it doesn’t necessarily tell you what happens when the whole country is exposed at once. This is because the economy-wide effects are absorbed into the intercept in the first analysis: they are constant across places in the cross-sectional comparison.

They cite as a prominent example the claim from a 2013 paper that “import competition explains one-quarter of the contemporaneous aggregate decline in US manufacturing employment.” The authors estimated the regional effect of increased Chinese import competition, then multiplied that coefficient by the national increase in Chinese import penetration.

…
Read more ⟶

NFTs were never about the tech

And the most successful NFT project proves it


Remember NFTs? Yeah, the premise was unusually empty even by the standards of over-hyped and under-utilized technologies, but let’s try to take it seriously as a technology for a moment.

The basic idea behind the Non-Fungible Token (NFT) is that while digital files are infinitely copyable, the NFT provides a way to establish ownership over a unique token associated with a particular file. The blockchain records who owns the token and any subsequent transfers. The token itself doesn’t actually contain the file; it (usually) just points to metadata that points to a file stored somewhere else. The blockchain is inefficient enough without trying to load it up with JPEGs. The tokens are non-fungible because each token is distinct; ownership of one implies nothing about ownership of any other.

So far, so good. The primary type of file people ended up associating with NFTs was ugly, randomly generated images. The most famous collection was the Bored Ape Yacht Club by Yuga Labs, a particularly grotesque series of cartoon hominoids dressed up in various costumes.

A weakness of NFTs is that they don’t inherently identify the file they supposedly establish ownership over, beyond “here’s where to find it”. There are various ways to deal with this. The Bored Ape Yacht Club used cryptographic hashes. Each of the 10,000 ape images was hashed, the hashes were put in a particular order, and the whole thing was hashed again, with the final result recorded on the blockchain. Anyone could therefore download the apes later, perform the same calculation, and confirm that these were in fact the same ugly hominoids, in the same order, that the creators had committed to at the beginning.

…
Read more ⟶