Full transcript
The Software Nobody Owns
0:00Right now, a radiology system at a
0:01hospital is deciding which scan to show
0:04a doctor first.
0:07A trading algorithm is processing tens
0:10of thousands of transactions per second.
0:14And both of these depend on software
0:16that they couldn't replace.
0:20Maintained by the same
0:23nobody.
Codd, Stonebreaker, and the Birth of Ingres
0:34In 1970, a computer scientist at IBM
0:37named Edgar F. Codd published a short
0:40paper.
0:42It described something called the
0:43relational model, a way of organizing
0:46data into tables, rows, and
0:49relationships. And this could be queried
0:51with mathematical precision. For 3
0:54years, nobody actually implemented this.
0:57That is, until two professors, one named
1:00Michael Stonebraker and Eugene Wong,
1:02decided to prove that the theory could
1:05work in practice. And they called their
1:07project Ingres.
1:11Surprisingly, it worked. Better than
1:13anyone actually expected. Ingres became
1:16the proof that relational databases were
1:19only just a theory. They were the future
1:22of computing.
1:24University shared the code. A generation
1:26of database engineers trained on it. The
1:28relational model went from just a simple
1:30academic paper to industry standard. And
1:34Ingres was a big reason for that. But
1:37then in 1980, Stonebraker did what
1:39researchers do when the research works.
1:42He started a company. And this wasn't
1:44the first time. In fact, Stonebraker
1:47would do this probably nine times across
1:49his career. Build the proof,
1:51commercialize it, and then move on to
1:53the next idea.
1:55Research that could be turned into
1:57product of some sort just simply wasn't
2:00done.
2:01But by 1985, he was back at Berkeley.
2:04Relational Technology ran without him.
2:07He'd spend a decade watching users hit
2:10the same walls with it.
2:13The database was simply built around a
2:15specific kind of work: payroll,
2:17inventory, accounting.
2:19Outside of numbers, text, and dates, it
2:22couldn't really store much else. And if
2:25you were a chemist with a molecular
2:26structure, or a mapmaker drawing a
2:29coastline, or an engineer designing a
2:32circuit, you couldn't just put your data
2:34into this database and call it a day.
2:36You had to break it into rows and
2:38numbers and rebuild it yourself every
2:40time you needed to use it for your
2:42project.
2:45It was the limit of his own creation.
2:48Most engineers who built something
2:49successful spend the next decade
2:52defending it. Stonebraker looked at his
2:54own success and saw what it couldn't do
2:56before he saw what it could. So, he
2:59thought, "What would I do differently?"
The Postgres Rewrite
3:01And that question became a brand new
3:03project, a complete rewrite, and he
3:06called it
3:07Postgres.
3:09Post-Ingres, named after the project it
3:12was leaving behind. But this time, it
3:14wasn't going to store just the world of
3:16business. It was going to be able to
3:18store everything. And so, funding came
3:20from DARPA, the agency that funded the
3:22internet, basically, from the Army
3:24Research Office, the National Science
3:26Foundation, a defense contractor called
3:28ESL. We have military public money going
3:32in to build what would become one of the
3:34most open pieces of software in history.
3:36But this ambitious idea came at one of
3:39the most hostile markets in tech
3:41history.
3:42In the database wars of the 1980s,
3:45performance was the only argument that
3:47really mattered. It was the only thing
3:50that translated into profit,
3:51essentially, especially when you're
3:53dealing with businesses. Fast reads,
3:54fast writes,
3:56fast everything. So, when Stonebraker
3:58said he wanted to build a database that
4:01was correct over fast, it was definitely
4:04an uphill battle. His bet was
4:06extensibility, custom data types, custom
4:09index methods, custom query operators.
4:12A database that wouldn't be a fixed
4:14product, but a platform that could grow
4:16as computing also grew, which was fast.
4:19Nobody in 1986 knew what that would
4:21mean, not even Stonebraker himself did.
4:24By 1992, the research was done with
4:27version 4.2, the final Berkeley release
4:30of Postgres. And after that, Stonebraker
4:33did what Stonebraker does.
4:36He started a company to commercialize
4:38Postgres, and venture capital funded it.
4:41Stonebraker left Berkeley for the second
4:43time.
4:49Two different theories about how
4:50software survives now began running in
4:52parallel, both with completely different
4:55ideologies on what makes software
4:58successful. One path, commercialization,
5:00resources, customers, direction, and the
5:03other, an FTP server at a university and
5:06whoever happened to
5:08find the address by coincidence.
Sponsor — Supabase
5:12This video is sponsored by Supabase,
5:14which for a documentary about Postgres
5:16makes a lot of sense. Supabase is built
5:18on that second path, the one that nobody
5:20bet on. When you spin up a Supabase
5:21project, what's running underneath is
5:23real Postgres, the same database this
5:25whole video's about, the one behind the
5:27hospitals and the banks. Supabase runs
5:29it as a managed dedicated database, and
5:31build the rest of its back end around it
5:33like authentication, file storage, and
5:35auto-generated APIs. You get all of that
5:37hosted without running the server
5:38yourself. And because it's standard
5:40Postgres, the whole ecosystem comes with
5:42it. Extensions like vector search for
5:44AI, PostGIS for mapping, and more. The
5:47standard tooling too. You can connect
5:49with PostgreSQL, PG dump or any
5:51PostgreSQL client that you already use.
5:53It's the same PostgreSQL that you run
5:55everywhere. And Superbase also puts work
5:57back into PostgreSQL itself. They
5:59acquired Oriole DB, a new storage engine
6:01from PostgreSQL, built by Alexander
6:04Korotkov, a PostgreSQL committer who now
6:06works at Superbase. And they've said
6:08that they wanted to upstream that work
6:10into core PostgreSQL. They use what
6:11these strangers built and give back to
6:13it. You can start your project for free
6:14now. Thanks again to Superbase for
6:16sponsoring today's video.
Postgres 95 and the Berkeley Handoff
6:18All right, back to it. Andrew Yu and
6:20Jolly Chen were PhD students at Berkeley
6:22in the mid-90s, finishing their
6:24education and prepping to move to
6:26industry.
6:27They were not database legends or men on
6:30a mission. They just simply needed to
6:32use PostgreSQL to behave like a real
6:34database, one that other systems could
6:36talk to. But one issue is that
6:37PostgreSQL used its own query language
6:40called Quel. Well, nobody outside of
6:42Berkeley had learned it. The rest of the
6:45database world had converged on
6:47something called structural query
6:49language or SQL for short. And
6:51PostgreSQL just didn't speak it.
6:55Any developer trying to connect it to an
6:57existing system hit a wall. The database
6:59might as well just not exist at that
7:01point.
7:03So, they did what programmers did. They
7:05added SQL themselves.
7:09Without a grant or a mandate, without
7:12asking Stonebraker himself, who was
7:13obviously busy building Illustra. They
7:16needed it for themselves. Putting it on
7:18the internet cost them nothing. Nobody
7:20else was going to do it. So, they did
7:22it. And they called it
7:24PostgreSQL 95.
7:28Just a version name, not really a
7:29project name. Someone else can take it
7:31from here. So, on September 5th, 1995,
7:33they sent the announcement from a
7:34personal email domain, that kind of
7:37distribution channel that didn't exist
7:39for software that anyone took seriously,
7:41even today to be honest.
7:43And
7:46they heard nothing back.
7:53You finished his education and joined HP
7:56while Chen also finished his and moved
7:57to the industry.
7:59The code went into the world without the
8:01authors.
8:02This time for the second time, the
8:03researcher built it and left.
8:06And then the students fixed it and left.
8:09Each handoff happened almost by
8:11accident.
The Volunteers Who Saved It
8:24Mark Fournier ran hub.org, a small
8:27Canadian hosting company operating out
8:29of a leased rack space in an era when
8:32the internet was run by people with
8:35large phone bills and hardware that
8:37they'd already bought. He'd been
8:38following the project from outside.
8:41And he thought it was worth preserving.
8:44He had what it needed, a server
8:47essentially.
8:49And so he offered it. The first
8:51non-university server in PostgreSQL's
8:53history given by someone the project had
8:56never asked.
8:58Bruce Momjian was a consultant in
9:00Philadelphia who needed an SQL database
9:02for his home Unix machine.
9:06And so he found Postgres 95, free with
9:08code legible enough to read.
9:11He started filing patches. No employer
9:13in the picture or no grant, just a
9:15mailing list where strangers in
9:16different time zones would review each
9:18other's work and either merge it or push
9:20back.
9:23And for years through the long stretch
9:25where the project had no obvious future,
9:28he kept coming back to it and
9:30surprisingly so did the others. And so
9:33by 1998, three years of volunteer
9:35patches, each one solving problems and
9:38sometimes creating another, had piled
9:41into a debt the project had not had the
9:43people to pay. So, for several months,
9:46it was unclear whether the project would
9:48stabilize or just completely collapse
9:50under its own weight, as open source
9:52software often does.
10:00In Pittsburgh, Tom Lane was running a
10:02small trading models business. He needed
10:04a database, like most developers too,
10:06and he chose Postgres for one reason and
10:09one reason only. The code was clean and
10:11readable. He filed a few patches, then a
10:14few more.
10:16He started looking at the query
10:17optimizer, the layer that decides how
10:19every single query runs, and the
10:21component most likely to produce the
10:24silently wrong answers before anyone
10:26caught them. And after two or three
10:27years, in his own words, he realized he
10:29was basically shrinking SSS work
10:32completely and spending all my time on
10:35Postgres stuff.
10:38The trading models faded into the
10:39background. The optimizer became his
10:42life's work almost by accident. Then
10:44Vada Makeev was contributing from 3,000
10:47km east of Moscow.
10:49He was working on MVCC, multi-version
10:52concurrency control, the feature that
10:54would define what Postgres felt like to
10:56use by letting readers and writers stop
10:58blocking each other.
11:01And he did this from a country whose
11:03economy had collapsed just years
11:05earlier, where a stable software job was
11:07not a given, let alone unpaid work on a
11:11foreign project with no clear future.
11:14His patches arrived in North American
11:15inboxes before anyone had even woken up.
11:19And when Lane opened his email in the
11:20morning, Makeev had
11:23already been working. The project had
11:24built a development cycle that did not
11:26sleep.
11:31That's when PostgreSQL 6.5 shipped in
11:34June of 1999 with Mickey's MVCC at its
11:37core. A stabilization release, proof
11:40that the community could hold together
11:42under pressure, deliver a major feature,
11:44and not break the existing system in the
11:46process. And that was the most important
11:48thing they ever proved. But meanwhile,
11:51Ingres sold to Ask, then to Computer
11:55Associates, then wandered through
11:57corporate owners for decades.
12:00And today it's called Actian. It is used
12:02by almost nobody.
12:06Illustra was acquired by Informix in
12:071997 and absorbed by IBM in 2001. And
12:11whatever it had been, it was something
12:13else now. The open version was on a
12:16donated server in Canada, maintained by
12:18people who fixed their own bugs. And the
12:20next decade would test whether that was
12:22really enough.
12:30For the next 30 years, the volunteer
12:32track would face the kind of pressure
12:34that kills open source projects almost
12:36immediately. CouchDB faded. RethinkDB
12:40ran out of money and shut down. MongoDB
12:43pulled its license to fend off cloud
12:44providers. Elastic did the same. And
12:47Redis did the same. Half the relational
12:49systems of the 1990s and early 2000s,
12:52Firebird, MSQL, the original Sybase,
12:55Informix, survived today as footnotes,
12:57if at all. And so Postgres faced the
12:59same waves. And it had less than any of
13:02them. No CEO, no marketing budget, no
13:05defensive moats, just the mailing list,
13:07the maintainers, and whoever happened to
13:10show up that week.
13:13The first sign that the volunteer model
13:15was working came from places nobody
13:17expected. When Oracle acquired Sun
13:19Microsystems, the European Commission
13:21ran an antitrust review. And in its
13:24January 2010 clearance, the Commission
13:26named PostgreSQL by name as the credible
13:30alternative that could constrain
13:31Oracle's market power in the database
13:33industry. A regulatory body in Brussels
13:36had decided that a project with no owner
13:38was substantial enough to hold a
13:40multi-billion dollar company in check.
13:43Let that sink in for a second.
13:45A year and a half later, Apple shipped
13:47OS X Lion Server. MySQL had been
13:49removed. PostgreSQL was running in its
13:51place. No press release, no
13:54announcements. Just engineers inside
13:56Cupertino who had reached the same
13:58conclusions as the regulators.
14:01And so for the first time, the volunteer
14:03track was being treated as load-bearing
14:05infrastructure by the institutions that
14:07ran the modern internet, not just as an
14:10interesting experiment.
Surviving NoSQL and the Cloud Era
14:13And that recognition came
14:15at exactly the wrong moment.
14:22In 2009, a generation of developers
14:24decided that the relational model was
14:27over. NoSQL arrived. MongoDB raised
14:30hundreds of millions of dollars on the
14:32argument that the rigid table and row
14:34structure of databases like Postgres was
14:38a legacy idea. Fine for accounting in
14:411985, but useless for the unstructured
14:43data that powered the modern web
14:45application era. Every conference, blog
14:47post, job posting reinforces message.
14:51Relational was the past and open source
14:53databases that didn't that got abandoned
14:56by their developer communities. The
14:58mailing list thinned, the patches
15:00stopped, the projects didn't shut down,
15:02they just stopped being relevant, which
15:04is the same thing. But slower. But on
15:07the Postgres mailing list, the
15:09volunteers had an argument on their
15:11hands. Some maintainers felt JSON
15:13support, the document format that NoSQL
15:16had popularized, would blur what
15:18Postgres was meant to be in the first
15:20place. A relational database that
15:21pretends to be a document database, they
15:23said, is just a worse document database.
15:28And others said that refusing to build
15:30it was handing the future to MongoDB.
15:33The debate ran for months between people
15:36who had never met with no CEO to break
15:38the tie. There was only the mailing
15:40list, the arguments, the patches, and
15:41the rule that nothing shipped until
15:43enough of the maintainers were
15:44convinced. And from the outside, it
15:47looked like a process that should not be
15:49able to make hard decisions.
15:51From the inside, it had been making hard
15:53decisions for almost 15 years.
15:55PostgreSQL at 9.2 shipped in 2012 with
15:58native JSON support. Two years later,
16:00JSONB, a format that gave the
16:02flexibility of a document database with
16:04the guarantees of a relational one, and
16:06the extensibility design from 1986,
16:09Stonebraker's most academic-sounding
16:10decision, had just shown what it was
16:13for.
16:16Postgres didn't have to choose between
16:18relational and document. It simply could
16:20be both, and that was by design.
16:23NoSQL didn't kill it. The volunteers
16:25built the missing capability themselves,
16:27and then kept moving.
16:29And the next wave came from a direction
16:31nobody saw coming.
16:35In the mid-2020s, every major cloud
16:36provider started selling managed
16:38PostgreSQL. Amazon's RDS for Postgres,
16:41Google Cloud SQL, Azure Database for
16:43PostgreSQL. Trillion-dollar companies
16:45taking the work of unpaid volunteers,
16:48charging customers for it, and keeping
16:51the revenue. And other open-source
16:52projects watched this happen, and just
16:55completely broke. MongoDB changed its
16:57license in 2018, and Elastic changed its
17:00license in 2021, and Redis changed its
17:02license in 2024. The argument was the
17:05same in each case. Cloud providers were
17:07extracting value from open code without
17:09contributing back. The community was
17:11being mined. Postgres didn't change its
17:13license, though, and that somehow made
17:16the world of difference.
17:18The maintainers had a choice.
17:20They could lock the project down, pull
17:21the permissive license that had been
17:23there since the absolute beginning, and
17:25then switch to something that restricted
17:27commercial use, force the cloud
17:28providers to pay.
17:30Or they could keep going as they had
17:31since 1996.
17:34And so, they made the choice to keep
17:35going.
17:38And then, something unexpected happened.
17:40Because the license stayed permissive,
17:42the cloud providers and the new
17:44ecosystem of companies built around
17:45Postgres developed a stake in the
17:47project's health. They started hiring
17:49maintainers, paying them salaries to do
17:51what they had been doing for free for so
17:53many years, funding feature work that
17:55volunteers wouldn't have had time for
17:57otherwise.
18:01And the thing that looked like
18:02extraction turned into the strangest
18:04possible alliance. The biggest companies
18:06in software ended up sponsoring the
18:08people who maintained the database that
18:10their products depended on.
18:17And this is where the project could have
18:19died for real. The history of open
18:21source software is full of projects that
18:22have started as community efforts and
18:25ended as corporate ones. The pattern is
18:27gradual. A sponsor hires a few
18:29maintainers. The sponsor's priority
18:30becomes the road map, and eventually the
18:32sponsor is the project, and the
18:34community is a logo.
18:37Postgres has resisted that pattern for
18:39going on 20 years. The reason is
18:40structural. No single company funds most
18:44of the maintainers. The funding is
18:45fragmented competing sponsors that each
18:48have a stake in the project staying
18:50neutral. None of them can capture it
18:52without the others objecting. The
18:54mailing list still works the same way it
18:56did in 1998. The patches still get
18:59reviewed from different employers who
pgvector, AI, and the Question of Succession
19:00don't always agree. And it looks fragile
19:02from the outside, but it has been
19:04fragile in the exact same way for 30
19:06years. It somehow keeps working. In
19:09April of 2021, a year and a half before
19:11ChatGPT existed, a developer named
19:13Andrew Kane shipped an extension called
19:15PG vector. It added vector data types
19:17and similarity search to post grass.
19:20He didn't ask permission. He didn't need
19:22to. The extensibility design from 1986
19:25had made it possible. He saw it coming
19:27and he built it. And after the launch of
19:29chat GPT, the market for vector
19:31databases exploded. Venture capital
19:33funded a half a dozen new entrants, each
19:35one with a dedicated database for
19:37storing coordinates that machine
19:39learning models produce. Post grass
19:41already had one sitting in a public
19:43repository available to anyone for free,
19:45written by one person who built it 18
19:48months earlier. And so in 2023, post
19:50grass took over my sequel as the most
19:52used database among professional
19:54developers for the first time in the
19:55survey's history. The next year it
19:57stayed on top. And this was only
19:59possible because they showed up for
20:01nearly 30 years. But before any of that,
20:03there was a moment in 2015. The
20:05Association for Computing Machinery had
20:07announced the winner, Michael
20:08Stonebraker, for ingress for post grass.
20:11And he said the biggest impact of post
20:13grass by far had come from the two
20:15Berkeley students he affectionately
20:17called grumpy and sleepy. He set a pick
20:19up team of volunteers, none of whom had
20:22anything to do with him or Berkeley, had
20:24been shepherding the open source system
20:26since 1995. And then he said this, "It
20:29is open source at its best. But I want
20:31to just mention that I have nothing to
20:33do with that. And that collection of
20:35folks, we all owe a huge debt of
20:36gratitude because they have robusted
20:38that code line and made it so it really
20:40works. And he didn't know all of their
20:42names, but he knew that they existed,
20:43knew that they had taken and built and
20:45made it into something that he couldn't
20:47have made." Standing at the podium
20:49receiving the highest prize in this
20:50field of work that included post grass,
20:53he said it. The thing that won was not
20:56what he built.
20:59So, who owns the database running your
21:01hospital, your bank, your airlines?
21:06Well, nobody. A professor who knew what
21:08an abandoned project looked like and
21:10moved on anyway. Two grad students who
21:13fixed a problem they needed fixed and
21:14then got on with their lives. A Canadian
21:17with a spare rack space. A consultant in
21:19Philadelphia who kept giving up evenings
21:21and weekends to a project that he had no
21:22commercial future. And dozens and dozens
21:25of volunteers over 20 years. Some of
21:27them now draw salaries from companies
21:30that depend on what they built, but none
21:32of them own it. And the crazy part is
21:34the volunteers who carried Postgres
21:35aren't young anymore. The same pick up
21:38teams don't break or thank to 2015 has
21:40been the same team more or less for two
21:42decades. Some are still in their day
21:44jobs. Some are paid by companies that
21:46depend on the project. Some have stepped
21:48back.
21:50Stonebraker had a successor. He didn't
21:52choose them. He didn't even know their
21:53names or didn't ask them to show up.
21:55They just simply did. And the
21:57maintainers don't have one. The
21:59generation of database engineers being
22:00trained today learn on cloud platforms
22:03where the work of running a database has
22:04been abstracted away from them. The most
22:06critical database infrastructure on the
22:08internet has no single owner. It belongs
22:11to whoever decides to show up. But the
22:13question is
22:15whether anyone actually will.