Nobody Is Driving

AI researchers are quitting, labs are talking about slowing down, governments are worried about China, and everyone has an incentive to keep accelerating anyway. Maybe we have been looking at the AI alignment problem one layer too low.

Share
Nobody Is Driving

We keep asking whether we can control superintelligence. The more immediate question might be whether we can control the race to build it.

"The oldest and strongest emotion of mankind is fear, and the oldest and strongest kind of fear is fear of the unknown."
— H.P. Lovecraft

Yesterday I was on a train, passing through Zurich on the way back from a peaceful vacation in the Austrian Alps. And yes, VCs need vacations too.

 I was going to write about something totally different this week, but I had my blissful, alpine peace disrupted by some guy talking about his artificial intelligence and robotics startup, conducting an obnoxious job interview for over an hour.

I can’t pin down exactly what he was saying; it all blended together in a sort of jargon soup. 

Something about fancy models, robots, humanoids, the future. All perfectly reasonable subjects, particularly for someone like myself who spends an unreasonable and unhealthy percentage of his waking life around technology.

And yet my immediate internal reaction was something close to:

“Please, for the love of God, talk about something else.”

Football. Divorce proceedings. The intricacies of Swiss recycling systems. A nasty rash that won’t go away. Your complicated relationship dynamics with your disapproving father.

Anything would be better. Really.

I realized, with a hint of guilt, that I have officially become exhausted by anything adjacent to artificial intelligence.

Which obviously presents something of a professional problem for a guy expected to study up on the latest in frontier tech to better understand and invest in it…

I invest in technology. I advise technology companies. I host a podcast where we routinely discuss artificial intelligence. I use AI quite frequently. I write about it often enough that an algorithm could probably diagnose me with Stockholm syndrome.

Hell, I’d love to start a fund that invests in anything that has nothing at all to do with this infernal technology.

And still, hearing another stranger discuss AI on public transportation made my eyes begin the long journey rolling toward the back of my tired skull.

AI has become weather at this point.

Ever-pervasive, it now shows up at work, in politics, in schools, in the stock market, in advertising, in warfare, in music, in dating, in customer service, in your browser, in your phone, and more and more, inside whatever previously undisturbed corner of life you had been hoping might remain devoid of any artificial intelligence.

After a quarter of an eternity listening to Mr. Roboto, my wife showed me a post on Instagram and asked me whether we were all indeed going to die before 2030.

I told her no.

"I think it's unlikely."

Which is still what I believe. I think.

Unfortunately, while saying this, a much stupider section of my brain had already begun constructing scenarios in which civilization had collapsed, and I was living somewhere deep in a rugged mountain wilderness, hunting reconnaissance drones for spare parts like a futuristic caveman.

In this vision, I was wearing scavenged computer wire around my neck. Silicon wafers dangled from my body like the teeth of animals killed by some grizzled prehistoric hunter. I had presumably developed a highly sophisticated barter economy involving batteries and radio transceivers.

Apparently my subconscious is working its way into Hollywood territory, having already developed a script and a several-season arc for a television adaptation. 

It’s 2026. And in this new, warped reality, my wife's question no longer sounds completely and utterly insane.

Everyone Has Become an Amateur Extinction Actuary

A week ago, a 27-year-old former Anthropic researcher named Jacob Coxon published his first post on X and announced that he had quit.

Coxon had worked at both Anthropic and OpenAI. He accused the companies of racing toward self-improving AI despite fears that the technology could eventually become a catastrophic, existential threat to our species.

Source: X

The post detonated, accumulating over 170 million views so far.

To make things more interesting, Coxon later told Axios that he had resigned roughly two months before his Anthropic equity would have vested.

Walking away from unvested equity is the kind of thing that definitely grabs my attention.

I am a venture capitalist, after all. I have seen grown adults tolerate extraordinary indignities like demonic bosses or painful company retreats in the presence of potentially lucrative stock options.

Well anyway, Coxon walked out the proverbial door.

His warning also managed to land at a pretty tense time, when more and more folks grow weary and fearful of whatever AI is with every passing day.

This summer, Pew found that 52 percent of Americans now say they are more concerned than excited about AI, compared with 37 percent in 2021. Seventy-one percent think AI will result in fewer jobs over the next twenty years. I’d be shocked if those numbers weren’t climbing higher as we speak.

The apocalypse debate has officially broken containment.

You no longer need to spend weekends on obscure rationalist forums discussing p(doom) with someone named NeuronLord420, or battling bots on Twitter.

Your spouse can now somewhat casually ask whether humanity survives the decade over breakfast.

The difficult part is figuring out what the hell you are supposed to tell them...

I decided to write this piece precisely because I’m not sure what the hell we’re supposed to respond to those types of questions anymore.

Who Exactly Am I Supposed to Believe?

As I took yet another swan dive down the rabbit hole this week, I noticed that almost immediately, people started poking holes in the Coxon story.

His account was new. He had barely posted publicly before. Then suddenly his first major appearance on the internet was a spectacular resignation manifesto about the possible extinction of humanity.

Some people wondered whether this was marketing. Or Chinese propaganda. Or some nefarious, little plot to stifle competition.

Other folks wondered whether his career prospects had actually improved enormously after becoming perhaps the most famous AI safety defector on our humble, little space rock.

Business Insider even asked whether quitting Anthropic might ultimately prove to be a pretty good career move. That’s a silly question to focus on right now, though.

It would also be one of the stranger marketing strategies ever conceived.

Imagine the meeting.

MARKETING EXECUTIVE: We need to increase awareness of Anthropic.
RESEARCHER: What if I publicly imply the thing we're building could kill everybody?
MARKETING EXECUTIVE: Interesting. Can we get that into a marketing carousel?

The problem is that I genuinely do not know. Neither do you. And right now, I’m suspicious of anyone who claims they do.

Coxon could indeed be exactly what he appears to be: a frightened researcher who decided the rapid ascent of frontier AI had finally crossed a personal line.

It’s entirely possible that this newfound public attention creates other opportunities for him in the future. Or maybe his concerns are sincere, AND his departure is still beneficial to his career.

Human motives sure are inconveniently capable of containing a multitude of ingredients.

The same problem pops up nearly everywhere else in this debate too.

Anthropic may sincerely believe highly capable AI could become dangerous in the not-too-distant future.

It is also hard to ignore that an expensive new safety regime would be much easier for Anthropic, OpenAI, or Google to survive than for some smaller lab trying to catch them.

Then the open-weight crowd makes the opposite case. Locking the world's most capable models inside a few giant corporations hardly sounds like a recipe for safety either. Fair enough. 

But once you release the weights, you don't exactly get to send everybody an email six months later saying "Sorry, turns out that one was more dangerous than we thought; please delete it."

Anthropic says it does not support banning open weights outright. OpenAI says regulation should target capabilities rather than smaller developers.

Those may be perfectly sincere positions.

They also happen to work pretty well for the companies already sitting at the frontier.

The people most qualified to warn us about frontier AI are also the very same people building frontier AI. And let’s not forget, the people best positioned to regulate it often depend on information produced by those frontier AI labs.

The investors investing trillions and trillions into the technology have financial reasons to believe it will become world-changing.

The politicians regulating it have economic and national-security reasons to make sure their own country stays ahead.

Even the jaded AI critics have ideologies and incentives of their own.

The open-source people think centralized control is dangerous, while the centralized labs think uncontrollable distribution is downright reckless.

We all have our receipts and blind spots and deep mistrust in others.

I wrote recently about how much of technology's growing public backlash ultimately comes down to trust. Tech keeps insisting that people simply need to understand the future better, while ordinary people increasingly wonder whether the people designing that future understand them at all.

This week managed to make that problem a hell of a lot worse.

We are being asked to personally price a potentially existential risk using information supplied by institutions whose incentives we cannot ever fully disentangle.

In the world of venture capital, uncertainty is essentially the inventory.

You look at incomplete information, questionable projections, strange founders, emerging markets, unknown competitors, and a future that refuses to sit still long enough to model properly.

Then you take a wild shot in the dark and make a decision anyway.

But I have never encountered another investment category where some of the people selling the asset are simultaneously warning that its tail risk may include the extinction of the buyer.

This all makes SaaS multiples feel refreshingly straightforward and transparent.

Unfortunately, the Bosses Are Also Worried

The Coxon story would be easier to dismiss if the companies involved responded by calling him a drama queen.

They did not. In fact, there has been a bona fide talent exodus lately coming from these labs.

OpenAI Chief Scientist Jakub Pachocki wrote earlier this month that he does not believe any laboratory has solved alignment and monitoring well enough to continue scaling "at maximum speed for much longer." He said he expects voluntary slowdowns and believes international coordination should become a top priority. I’m not holding my breath there.

OpenAI also announced that it had achieved its goal of building an automated research intern, a system capable of assisting with the very research used to build more advanced AI. How heartwarming.

Seems a bit unhinged to keep announcing these things in tandem with one another.

We need to slow down.
Also, we built a machine that helps us go faster.

Anthropic CEO Dario Amodei then published an essay titled We Must Pace the Frontier.

Source: Dario Amodei

Amodei argued that frontier AI development needs to proceed more slowly so safety practices can catch up. He proposed permanent third-party evaluators inside major labs, coordination between companies, and eventually international coordination…

Oh, and we also need to all hold hands and sing Kumbaya at the same time.

Sam Altman and several other prominent AI leaders expressed support for some version of slowing or pacing frontier development.

Even the language is revealing.

None of these people actually wants to "stop."

That sounds irresponsible. They want to "pace." Much nicer and neater.

Like a marathon runner checking his heart rate on his watch rather than a civilization wondering whether somebody should maybe unplug the experimental all-knowing god-machine until next Tuesday.

Mary Shelley arrived here before anyone else.

Frankenstein was subtitled The Modern Prometheus, and Victor eventually warns about "how dangerous is the acquirement of knowledge."

Source: Wikipedia

The unsettling part of this Sci-Fi horror classic is the creator's relationship with creation.

Ambition comes first, and then responsibility arrives much later to the finish line, looking slightly disheveled and out of breath.

We Should Slow Down

Great.

Problem solved.

Everyone agrees to slow down. Let’s do it. What are we even still talking about this for?

AI development proceeds carefully. Safety eventually catches up. Humanity finally receives those promised cancer cures and productivity gains and scientific breakthroughs and cheap intelligence, and perhaps a little friendly robot capable of folding a fitted sheet.

Unfortunately, another participant has entered this here meeting.

Competition.

Imagine two frontier labs. Let’s call them Misanthropic and ClosedAI.

Both believe continuing at maximum speed could become dangerous.

Both claim to prefer a world where every major lab proceeds more cautiously.

Neither wants to become the only company proceeding cautiously while its competitor keeps accelerating, pedal to the metal.

So someone says:

RESEARCHER: We need to slow down.
LAB: Absolutely.
RESEARCHER: Great. Let's slow down.
LAB: We will.
RESEARCHER: Wonderful.
LAB: As soon as they do.
OTHER LAB: What?

There is a famous game-theory debacle hiding inside this abject stupidity.

In a prisoner's dilemma, individuals pursuing perfectly rational self-interest can collectively produce an outcome worse for everyone.

AI development has now taken this exact shape.

  • Anthropic cannot simply pretend OpenAI does not exist.
  • OpenAI cannot pretend Google and Meta do not exist.
  • Meta cannot pretend Chinese open-weight AI labs don’t exist.
  • Investors cannot pretend competitors are not allocating trillions to the same opportunity.
  • Employees cannot pretend somebody else will not slot in and take their job.

And every time one actor on this stage accelerates, everyone else receives another reason (or multiple) to do the same. 

This does not necessarily require a villain, which is a bit bothersome.

Villains are comparatively quite convenient.

You lock them up and throw away the key.

Systems of individually rational people producing collectively irrational outcomes are much harder to account for. 

In this particular case, people don’t need to have sinister intentions. Merely reacting to the moves of others seems sufficient enough for mayhem to ensue.

Then Somebody Says China

At this point, the corporate prisoner's dilemma gets promoted to international relations.

Source: Unhinged Psychopath Guy's Twitter

Trump has rejected calls for an American AI pause, arguing that slowing the United States could allow China to gain an advantage. House Speaker Mike Johnson made essentially the same argument this week, saying a moratorium could surrender America's competitive advantage and create national-security risks.

There it is. The convenient boogeyman. The ultimate conversation-ending word in American technology policy:

China.

Source: South China Morning Post

China, unsurprisingly, views American AI policy through its own strategic lens. Chinese commentary has attacked slowdown proposals as potentially serving American technological containment, while the United States continues restricting access to its advanced chips and other critical technology.

And now the riptide loop is much harder to escape.

Anthropic worries about OpenAI. American policymakers worry about China. China notices that America is restricting its access to the technology required to compete. Every move becomes evidence for the next move. And the next move. And the next dozen moves and so on.

Nobody even has to be lying.

America can genuinely believe it would be dangerous to surrender leadership in a technology this important. Washington does not have to be cynical about this. Falling behind China in a technology this consequential can plausibly look like a national-security failure.

China can genuinely interpret American restrictions as an attempt to stop it catching up. AI labs can genuinely worry about catastrophic risk while also believing that allowing a less cautious competitor to get there first would be even worse.

Every explanation makes sense on an island.

Together they form something remarkably stupid.

Last week I wrote about autonomous weapons and a similar dynamic. Humanity would very much like to keep humans meaningfully involved in lethal decisions, but warfare rewards whichever side can process information and act faster. 

The AI race has the same ugly characteristic. Everyone can prefer restraint in theory while facing incentives that punish restraint in practice.

It’s amusing to try and apply that logic to literally any other industry.

Suppose Airbus unveiled a revolutionary new passenger aircraft. The plane is hailed as extraordinary. And faster. Cheaper too. Oh, and more efficient. It’s capable of transforming all of global transportation forever, but there’s just one tiny issue…

Engineers haven’t yet completely solved a failure mode that potentially causes the flight-control systems to behave erratically.

Nobody knows the probability. It could be tiny or insignificant. All we know is that there isn’t enough information yet.

The chief engineer steps up to the microphone.

"Our new aircraft is unlike anything ever constructed. We believe it could transform human transportation. Our engineers also think there's a meaningful chance it becomes uncontrollable and kills everybody aboard. Unfortunately, Boeing is building one too, so we have to start selling tickets."

Someone asks whether perhaps the plane should remain on the ground.

"We would love to."
"Great."
"Unfortunately Boeing is developing one too."

Then someone from Washington arrives to explain that China has a prototype.

Tickets go on sale Tuesday.

AI is obviously not commercial aviation. The AI sphere operates under its own rules.

Source: ACRATS

We built an elaborate system around aviation that is capable of telling somebody no, even when enormous amounts of money and competitive advantage are involved. A catastrophic unresolved failure mode can keep a plane on the ground.

The Machinery Around the Machine

With AI, competition itself is increasingly becoming the explanation for why nobody can remain on the ground.

Which brings me back to the thing that has been haunting me about alignment.

Most AI alignment talk is about making sure an advanced machine does what humans actually want it to do, rather than interpreting “solve climate change” as an instruction to remove everything with a carbon footprint.

Worthwhile problem.

I would personally prefer not to end up as raw material for whatever deeply stupid objective the superintelligence misunderstood.

But look at the machinery already surrounding it.

Labs, researchers, venture funds, public markets, Nvidia, sovereign wealth funds, chip restrictions, data centers, Washington, Beijing, Brussels convening some committee with “responsible frontier framework” buried somewhere in the title.

Nobody designed the whole arrangement, and nobody is really in charge of it, steering the ship.

And somehow the whole contraption keeps arriving at the same answer: Go faster.

Build the bigger cluster or raise the next round. Hire the legendary sought-after researcher before someone else does. Secure the chips and train the model. Open a hyperscale data center bigger than Rhode Island. 

Do not be the absolute moron who slowed down first.

Due to my choice of career, I am professionally enthusiastic about a massive chunk of this.

Competition is useful. Capital finds opportunity. Startups make incumbents uncomfortable. People attempting to murder one another commercially have produced everything from cheaper televisions to better cancer treatments.

But markets are responsive, not omniscient or wise.

Get enough incentives pointing in one direction, and they are extremely good at making a lot of people run that way.

They are much less concerned, however, with the fact that there may be a gigantic cliff awaiting them.

The AI industry spends an enormous amount of time worrying about a future machine relentlessly optimizing toward some badly specified objective.

Meanwhile, look at us…

Anthropic cannot ignore OpenAI. OpenAI cannot ignore Google or Meta. Washington cannot ignore Beijing. Investors see other investors writing checks. Researchers see other researchers joining frontier labs.

Nobody wants to be the only pacifist at an arms convention.

Each decision makes sense from roughly six inches away. 

That is what bothers me more than any individual CEO or researcher.

You can get somewhere very strange without anybody deciding to go there.

I am also implicated in this. I invest in companies. I like technological progress. I make money, in theory at least, when new technologies become important. My incentives are sitting on the same table out in the open, just like everybody else's.

Maybe that is why the sudden certainty surrounding AI annoys me so much.

Everyone seems awfully certain. The doomers, the accelerationists, the labs, the open-weight crowd. The guy on Reddit who has spent three months replying ‘you fundamentally don't understand transformers’ to strangers definitely knows.

I don't. I still don't think anybody does…

Maybe the people predicting catastrophe have dramatically overestimated how quickly intelligence scales. 

Or maybe alignment turns out to be difficult but manageable. 

Or maybe today's models are much further from anything resembling autonomous superintelligence than the most frightened researchers believe.

I hope so.

There is already plenty to worry about without Skynet: jobs, surveillance, fraud, propaganda, cybersecurity, political power, the concentration of machine intelligence inside a handful of companies, and the possibility that we outsource enough thinking that our own brains become mostly decorative mashed potatoes.

Then the people actually building frontier AI started publicly debating whether everybody should ease off the throttle.

How amusing.

My wife asked me whether we are all going to die before 2030.

I still think the answer is probably no.

But I cannot give her a useful probability, and I have grown suspicious of anyone who can.

There is no historical dataset or precedent for this. There is no agreed model. The experts disagree wildly, and almost every source of information has some incentive or bias attached to it.

If someone tried to value a startup to six decimal places under those conditions, I would assume they had suffered a grievous head injury.

With AI extinction, somebody puts the number on a graph.

And while everyone argues about whether the correct probability is 0.1 percent or 30 percent, the machinery keeps chugging along.

We tend to imagine that if AI becomes dangerous, there will be a moment when humanity chooses to cross some obvious line. An alarm sounds. A model escapes. An exhausted researcher runs into a control room screaming that it has become self-aware. Sam Altman reaches for a large red button that inexplicably and perhaps cartoonishly says AGI.

Arnold Schwarzenegger materializes naked behind a Zurich tram stop.

At least that would be clear. And pretty epic.

Reality is rarely that courteous. There may be no singular, obvious decision where humanity chooses to proceed.

There may just be another model because the previous model existed. Another billion dollars because somebody else just raised two in a Series G. Another cluster because China wanted to build one. Another export restriction because the cluster got built. Another government subsidy because the export restriction worked. Another safety concession designed carefully enough that nobody has to actually stop.

You can travel a very long way before anyone notices nobody chose the destination.

The strange part is that the race itself already seems harder to control than any one company or government involved in it. Or the technology we’re currently fighting over.

Nobody gave that race a constitution. Nobody can unilaterally redirect it. It merely reacts to incentives, absorbs new participants, and punishes anyone who stops while the others keep moving.

And its objective seems remarkably simple:

Do not fall behind.

That does not mean it ends badly.

It may produce astonishing things. New medicines. Better science. Abundance. Technologies I would gladly invest in and use.

It’s certainly a version of the future I can get behind.

I just no longer find “everyone has excellent reasons to keep accelerating” particularly reassuring. Not in the least.

Everyone keeps talking about whether we can control superintelligence. We have not yet demonstrated that we can control ourselves.