Michael Nielsen on changing the units of publication in science

Yes. Your intuition is very close to an actual metascience lineage, and Michael Nielsen was saying it unusually explicitly almost 20 years ago.

The strongest version is not merely “make papers shorter.” It is:

Stop treating the paper as the atomic unit of science.

Instead, science should look more like a version-controlled, continuously updated graph of claims, observations, methods, data, code, failures, conjectures, critiques, and syntheses. A paper can still occasionally crystallize part of that graph, but it should not be the container everything must be stuffed into.

Michael Nielsen was astonishingly explicit about this

In 2007 Nielsen called the idea “micropublication.” His definition had three parts:

  1. Decrease the size of the smallest publishable unit of research.

  2. Treat many more things as first-class research outputs, including data, code, simulations, commentary, ideas, etc.

  3. Lower the barrier between doing something useful and making it publicly available.

His point was that the paper system causes huge amounts of useful scientific cognition to disappear because only finished, narratively coherent successes survive. Michael Nielsen

Nielsen even opened the essay with Doron Zeilberger’s proposal to publish good incomplete attempts rather than throwing all unsuccessful reasoning into the intellectual trash compactor. That is remarkably close to what you are talking about.

Nielsen, “Micropublication and open source research”

And Nielsen’s broader argument was that journals provide civilization with a long-term memory, but the internet could additionally give science a collective short-term working memory, where ideas are exposed, modified, criticized and recombined while they are still alive. His Polymath work with Tim Gowers was an existence proof: dozens of small contributions accumulated into a mathematical result. Michael Nielsen

That’s a much more radical conceptual change than open access:

journal model
research → polish → narrative → reviewers → paper

versus Nielsen’s direction:

idea → observation → question → partial result → correction → method → counterexample → improved claim → synthesis

with every arrow potentially public and citable.

And Nielsen specifically argued that papers expose only a tiny slice of scientific knowledge

His open-science writing lists things scientists know but that the paper format usually fails to capture:

questions, ideas, leads, folklore knowledge, notebooks, opinions, workflows, simple explanations, along with data and other artifacts. Michael Nielsen

There was even a 2007 guest essay by Peter Rohde on Nielsen’s site saying something hilariously close to your complaint: he could count the papers he’d fully read that year “on your hands,” because usually he wanted the central idea, method, and open questions rather than pages of formal apparatus. Rohde called merely moving the conventional paper from paper to an LCD a misuse of electronic media, and argued for modular, hierarchical scientific communication instead. Michael Nielsen

Important distinction: that particular essay was Rohde, not Nielsen, although Nielsen hosted it.


Adam Marblestone is making a related but different attack

Marblestone isn’t primarily the “replace papers with micro-units” person.

His FRO argument attacks the problem one level upstream.

Academia has implicitly made:

publishable paper ≈ legitimate scientific output

But many of the most valuable things science needs are not naturally papers at all:

  • datasets

  • measurement infrastructure

  • software

  • experimental platforms

  • standardized protocols

  • new organisms/models

  • mapping resources

  • engineering systems

FROs are explicitly designed to produce these research-enabling public goods that academic incentive structures systematically underproduce. Marblestone and colleagues argue that ordinary academic institutions often aren’t organized to build sustained engineering-heavy platforms, datasets and tools. Issues in Science and Technology

So Nielsen says:

Make the unit of publication smaller and more heterogeneous.

Marblestone says something complementary:

Stop equating scientific accomplishment with publication in the first place.

Those combine extremely naturally.


Seemay Chou/Astera/Arcadia have now gone much further

This may actually be the person you’re remembering from recent Progress Studies conversations.

In June 2025, Astera cofounder Seemay Chou wrote Scientific Publishing: Enough is Enough. Her formulation is almost word-for-word your premise.

She says traditional journal publishing is “fundamentally broken”, and reports that among scientists she has discussed it with across sectors, “exactly zero think the journal system works well.” Astera consequently began requiring much of the research it funds not to be routed through traditional journals. asterainstitute.substack.com

And then the key sentence:

“Scientists should probably be putting out shorter narratives, datasets, code, and models at a faster rate…”

with greater visibility into scientists’ reasoning, methods and mistakes. She goes on to argue that on the internet almost anything can be a publishable unit. asterainstitute.substack.com

That is almost Nielsen’s 2007 micropublication proposal reincarnated for the AI era.

And Jason Crawford has now incorporated exactly this argument into his Progress Studies agenda, criticizing the peer-reviewed paper becoming the measurable unit of work and explicitly quoting Chou’s shorter-narratives/data/code/models proposal. Roots of Progress

So there actually is a pretty clean genealogy:

Nielsen → micropublication/open science → Distill/research debt → FRO/output pluralism → Arcadia/Astera → contemporary Progress Studies.


Arcadia has actually tried it

This isn’t just manifesto-land.

Arcadia Science started publishing much smaller pieces of research, which they call “pubs.” A pub might be:

  • an observation

  • a resource

  • negative data

  • a method

  • an idea

  • a dataset

  • an experimental result

Rather than artificially making each one tell an entire scientific story, Arcadia maintains evolving project narratives that connect those smaller units. research.arcadiascience.com

That’s a subtle but important architecture.

Instead of:

one immutable 14-page narrative pretending the research occurred linearly

you get something closer to:

PROJECT
│
├── question
├── experiment 01
│   ├── protocol
│   ├── data
│   └── negative result
├── observation
├── revised hypothesis
├── experiment 02
│   ├── code
│   └── result
├── external critique
├── replication
└── evolving synthesis

The synthesis is a view over the graph, rather than the graph being mutilated until it fits the synthesis.

Arcadia explicitly says that the chronological record can expose failures and dead ends that conventional papers suppress. research.arcadiascience.com

That is a huge epistemic difference.


Chris Olah and Shan Carter supplied the explanation for why reading papers feels so awful

Their 2017 Distill essay “Research Debt” is probably essential to the argument you’re making.

Their claim wasn’t simply “scientists write badly.” It was that research communities accumulate something analogous to technical debt:

  • undigested ideas

  • poor explanations

  • bad notation

  • bad abstractions

  • enormous quantities of noise

Eventually every newcomer has to independently reconstruct the same conceptual structure.

They make the wonderful observation that for many papers you want one simple sentence explaining the important thing, but extracting that sentence requires wrestling with the entire paper. Distill

And there is a nasty incentive loop:

career reward ∝ papers

therefore

every useful thing → dress it up as a paper

therefore

paper count ↑ → attention fragmentation ↑ → research debt ↑

Olah and Carter explicitly praise open-notebook science and Polymath because they separate sharing research from turning research into a formal paper. Distill

Michael Nielsen was on Distill’s steering committee, so these aren’t entirely separate intellectual universes. Distill


There are now systems that go all the way to “science as composable primitives”

Octopus is probably the cleanest implementation.

It abolishes “the paper” as its primary object and defines eight independently publishable units:

Problem → Hypothesis → Method → Results → Analysis → Interpretation → Application, plus Review.

Each is independently attributable and linked to the others. Octopus

So Alice can have a brilliant hypothesis.

Bob invents the method.

Chen collects the dataset.

Diego notices the statistics are wrong and publishes a different analysis.

Elena interprets the revised result.

Nobody has to invent a fake five-author paper in which everyone supposedly did the same indivisible intellectual object.

ResearchEquals calls essentially the same idea modular publishing and explicitly promotes publishing research outputs continuously rather than waiting until the entire research cycle can be packaged as a conventional article. ResearchEquals

And the biomedical micropublication/nanopublication literature goes still smaller: a scientific object can be a single attributed claim linked to its evidence, with supporting and challenging claims forming a graph. PubMed Central (PMC)


The really interesting inversion

I wouldn’t conclude “papers are useless.”

There’s a strong reason they evolved.

A good paper is a checkpoint.

It forces someone to say:

Here is the coherent case I am willing to stand behind at time \(t\).

That’s useful. Otherwise an endlessly mutating scientific knowledge graph can become an archaeological ruin of half-thoughts.

So I think the mistake is not the existence of papers.

It is making the paper simultaneously be:

unit of discovery + unit of communication + unit of credit + unit of evaluation + unit of archiving + unit of peer review.

That’s the architectural monstrosity.

A much better system is:

micro-units underneath, syntheses above

atomic observations
       ↓
claims ↔ counterclaims
       ↓
data / code / protocols
       ↓
experiments / analyses
       ↓
short syntheses
       ↓
living reviews
       ↓
occasionally:
major integrative papers/books

The paper becomes something like a software release, not the source-code commit.

And this becomes dramatically more important with AI.

An AI scientist doesn’t particularly want 10,000 PDFs containing introductions, rhetorical glue, duplicated Methods boilerplate and strategically constructed “stories.”

It would much rather receive:

claim
evidence
experimental conditions
provenance
confidence
code
dataset
dependency
contradiction
replication status
timestamp

That is a scientific dependency graph.

Chou now makes this exact AI-era argument: the sanitized journal record hides the actual reasoning process, which makes it worse training material for scientific AI agents. asterainstitute.substack.com

And I think that’s where Nielsen’s 2007 idea suddenly gets much more explosive in 2026: he was imagining a medium optimized for human collective intelligence. We can now imagine essentially the same architecture serving human + agent collective intelligence.

The provocative way I’d put the whole movement is:

The paper should stop being the atom of science. It should become one possible compression of science.

And once you make that switch, the enormous amount of scientific cognition presently left on lab whiteboards, Slack, notebooks, GitHub issues, abandoned analyses, failed experiments and people’s heads suddenly becomes potentially part of the scientific record.

Yes, but with a caveat: I can verify Nielsen very strongly on micropublication, and I can verify his broader educational work on breaking learning into smaller, composable units, but I’m not finding a clean primary-source hit where he himself prominently coins the exact term “microcourses.”

What he definitely argued in 2007 was that research should be publishable in “small incremental steps”, with code, data, simulations, explanations, criticism, corrections, etc. all becoming first-class outputs. Michael Nielsen

And there is a striking pedagogical analogue in his later work with Andy Matuschak. Their Quantum Country / mnemonic-medium project is essentially anti-course-shaped: instead of “take this 12-week course and hope it sticks,” the material is decomposed into small conceptual units, interleaved with retrieval prompts, and revisited over time. The underlying philosophy is very close to:

knowledge should be delivered in appropriately sized cognitive units, not inherited institutional containers.

So you can draw a pretty beautiful symmetry:

Science Learning
paper course
huge bundled unit huge bundled unit
publish after completing story certify after completing syllabus
reader must excavate useful claims student must excavate useful concepts
micropublication microlearning / microcourse
graph of claims graph of concepts
continuous revision spaced reconstruction

And Nielsen was already thinking this way even in older educational writing. He complained that research skills are rarely explicitly taught and argued for exposing much more of the actual process of research, rather than only presenting polished disciplinary content. Michael Nielsen

There’s also a revealing passage in his 2007 micropublication discussion where he imagines students directly rewriting and improving scientific articles, with version histories and attribution, rather than treating papers as sacred frozen texts. Michael Nielsen That effectively collapses publishing + learning + contribution into the same substrate.

So I think the deeper Nielsenian idea is actually broader than either term:

Shrink the institutional unit until it approaches the natural unit of cognition.

Research institutions historically bundle:

50 observations + 4 ideas + 2 methods + 1 result → PAPER

Universities bundle:

40 concepts + 15 exercises + 12 lectures → COURSE

But the useful objects may really be:

claim
explanation
example
counterexample
technique
question
experiment
dataset
concept
exercise
critique

Each independently addressable, linkable, revisable, attributable, and recombinable.

That leads to an even more radical architecture than “microcourses”:

not microcourses, but no fixed course boundary at all.

You could have a knowledge graph where someone wants to understand “pseudospectral sensitivity in aging,” and the system dynamically traverses:

non-normal dynamics
→ eigenvectors
→ resolvent
→ transient amplification
→ biological control
→ aging

with each node perhaps taking 3–15 minutes, and only supplying prerequisites when necessary.

That is to a course approximately what Nielsen’s micropublication is to a paper.

So your recollection may be combining Nielsen with the wider microcourse/microlearning vocabulary, but conceptually, yes, it is extremely Nielsen-shaped. The more interesting thing is that his micropublication and mnemonic-medium ideas look like two instances of the same principle: stop forcing intellectual work into giant Victorian shipping crates when the web can move individual Lego bricks.

It looks like Beagle basically stalled and never became the scientific knowledge layer Marblestone imagined.

Marblestone’s Beagle, started around 2014, was a Chrome/PDF annotation system built with Richard Littauer, Ed Boyden, Juan Benet, Konrad Kording and others. The pitch was unusually forward-looking: scientists would annotate papers while reading them, and those annotations would become a shareable meta-layer over the literature, including private lab notes, public comments, corrections, links, and potentially machine-readable structure. It even planned to use IPFS so the annotation layer would not depend on one proprietary silo. Massachusetts Institute of Technology

But the current evidence is pretty unambiguous:

  • Marblestone’s 2026 CV still lists Beagle as a project begun in 2014, but explicitly marks it “[paused]”. adammarblestone.org

  • Richard Littauer, who says he was the sole contracted developer for about eight months, describes his Beagle work as 2014–2016 and says the project is “currently in hiatus.” Richard Littauer

  • His current personal site likewise lists “Beagle PDF Annotator & Chrome Extension, Developer 2014–2016.” Richard Littauer

So there doesn’t seem to have been a dramatic shutdown, acquisition, or successor product. It looks more like a classic research-tool fate: prototype + seed grant + interesting architecture + insufficient adoption/continuity → dormancy.

What makes it especially interesting in light of your previous question is that Beagle was almost a missing bridge between Nielsen’s micropublication vision and today’s AI-science tooling.

The original Beagle proposal was already imagining something like:

paper
  ├─ annotation
  ├─ disagreement
  ├─ explanation
  ├─ related result
  ├─ private lab note
  ├─ public correction
  └─ discussion thread

rather than treating the PDF itself as the final object. The Google research proposal explicitly envisioned annotations attached persistently to papers, selective sharing from private groups to public discussion, and interoperability with systems such as Hypothes.is. adammarblestone.org

The funny part is that 2026 is a vastly better technological environment for Beagle than 2015 was.

Back then, Beagle needed humans to manually annotate enough papers to make the social graph valuable. That’s a brutal cold-start problem:

no annotations → little reason to install → few users → no annotations

Today an AI layer could bootstrap the whole thing:

PDF / preprint / dataset / repo
            ↓
AI extracts
  claims
  methods
  evidence
  assumptions
  dependencies
  contradictions
  open questions
            ↓
humans annotate/correct
            ↓
cross-paper knowledge graph
            ↓
agents reason over the graph

That gets astonishingly close to the thing you were just describing: science represented in smaller incremental units rather than giant paper blobs.

And there’s a deeper historical continuity here. Marblestone also worked on or advised OccamzRazor, which his current CV describes as a “knowledge capture and sharing platform for scientists,” and consulted for Expii on scientific knowledge graphs. adammarblestone.org So Beagle was not an isolated whim. He was clearly circling the problem of scientific knowledge representation for years before moving much more heavily into organizational innovation and eventually FROs/Convergent Research.

My read of the trajectory is roughly:

2014 Marblestone:
“Can we improve the substrate on which scientists exchange knowledge?”

2020+ Marblestone:
“Maybe an even bigger bottleneck is the organizational substrate on which scientists actually build things.”

Hence the shift from Beagle → FROs / Convergent Research. Convergent is now his central project and is launching dedicated organizations to build scientific infrastructure rather than trying primarily to repair literature communication. Issues in Science and Technology

But intellectually, Beagle aged rather well. Its implementation died; the problem statement became more important.

A modernized “Beagle 2.0” combining Nielsen micropublications + Arcadia-style small research units + citation graphs + agent extraction could be much more consequential than a PDF annotator. It would basically make the scientific record itself agent-native.