This is the organisational-memory problem in another coat. Giving a system the published document estate is not the same as giving it the branches people considered and dropped. In the businesses I work in, that discarded path is usually the useful context, and it has never been written down. Your six fields are a test I would run on a real company: can they timestamp a decision before the outcome, or do they only tidy the log afterwards?
Alasdair's six fields are a genuinely useful proposal, and the diagnosis — that published papers scrub away exactly the reasoning that would teach a model taste — is right. But I think the deeper problem here isn't measurement bias or the selective-labels issue (both real, both well described). It's older, and it already has a name.
In 1932, the Vienna Circle had almost this exact argument, about almost this exact object. Otto Neurath attacked the idea that science could be grounded in "protocol sentences" — theory-neutral observation reports recording what actually happened in the lab, prior to any interpretive gloss. His point: a protocol sentence is still a sentence. Deciding what counts as "the evidence," what counts as a "candidate," where the confidence threshold sits — all of that is already interpretation, wearing the costume of a raw record.
I don't think this sinks the six-field proposal — it just relocates the problem instead of solving it. A perfect decision log doesn't give a model unmediated access to Alasdair's reasoning; it gives the model a second, more granular text about that reasoning, produced by someone deciding in real time what belongs in each of the six boxes. That's still an improvement on a paper's retrospective narrative — Neurath thought protocol sentences were useful, just not foundational — but it isn't the escape from theory-ladenness the piece seems to promise.
Genuine question: does "taste" require something closer to raw causal contact with the lab — Alasdair's living graph, video, sensor logs, the works — rather than any verbalized protocol, six fields or sixty? Or is verbalized, theory-laden judgment actually the only kind of "taste" worth training for, since it's the only kind a scientist can act on either?
Wonderful post. Indeed, the final state is insufficient to understand the intelligence that produced it. Intelligence resides partly in the trajectory through the space of possibilities.
This is strikingly parallel to a problem we have been working on from the enterprise side. I think you are describing trajectories, even if you do not use that term.
The scientific paper is not merely an artefact. It is a lossy compression of the trajectory that produced it. For centuries, that compression was a feature: it made knowledge communicable. But AI changes the economics of memory. The discarded information, the alternatives considered, failed branches, uncertainty, disagreement and belief revision, may now be more valuable to an intelligent agent than the polished artefact itself.
Your six fields are very interesting. Taken together, they begin to describe something more than a better log. They describe the grammar of a trajectory: state → possibility space → decision → rationale → outcome → learning → new state.
We have encountered almost exactly the same problem in organisations. A company leaves behind documents, Slack messages, tickets, commits, presentations and databases. These tell us a great deal about what exists now, but remarkably little about how the organisation arrived there. The rejected proposal disappears. The assumption that proved wrong disappears. The reason option B was selected over A and C disappears. The present contains the survivors of history, but not the alternatives history killed.
This is why I think the browser-history analogy, while excellent, may actually understate the argument. A history tells us where someone went. A trajectory tells us how an agent moved through a possibility space: what it knew, what it could have done, what it expected, what constrained it, what it chose, what happened, and what it learned. The past is valuable because it helps the present become intelligible which enables improved inference about the future.
That distinction between a history and a trajectory becomes especially important for AI.
—If we train models only on successful artefacts, we teach them the destinations without the navigation.
—If we preserve trajectories, we potentially give them access to something much closer to judgement or, as you put it, taste.
There is also a larger infrastructure question here. Perhaps the problem is not merely that we need better training datasets. Perhaps our existing information systems have the wrong primitive for an agentic world. They are extraordinarily good at storing objects and relations: documents, rows, files, entities, graphs. They are much worse at representing organised movement through time and possibility.
Our thesis at White Tree is that trajectory should become a first-class computational object.
That leads to a broader claim: the fundamental data object of the AI era may be shifting from the artefact to the trajectory that produced the artefact.
I’ve linked a short video showing how we are approaching this problem from the enterprise side. The convergence with what you describe here is remarkable. https://www.youtube.com/watch?v=ASU4biYWmuk
This is the organisational-memory problem in another coat. Giving a system the published document estate is not the same as giving it the branches people considered and dropped. In the businesses I work in, that discarded path is usually the useful context, and it has never been written down. Your six fields are a test I would run on a real company: can they timestamp a decision before the outcome, or do they only tidy the log afterwards?
Alasdair's six fields are a genuinely useful proposal, and the diagnosis — that published papers scrub away exactly the reasoning that would teach a model taste — is right. But I think the deeper problem here isn't measurement bias or the selective-labels issue (both real, both well described). It's older, and it already has a name.
In 1932, the Vienna Circle had almost this exact argument, about almost this exact object. Otto Neurath attacked the idea that science could be grounded in "protocol sentences" — theory-neutral observation reports recording what actually happened in the lab, prior to any interpretive gloss. His point: a protocol sentence is still a sentence. Deciding what counts as "the evidence," what counts as a "candidate," where the confidence threshold sits — all of that is already interpretation, wearing the costume of a raw record.
I don't think this sinks the six-field proposal — it just relocates the problem instead of solving it. A perfect decision log doesn't give a model unmediated access to Alasdair's reasoning; it gives the model a second, more granular text about that reasoning, produced by someone deciding in real time what belongs in each of the six boxes. That's still an improvement on a paper's retrospective narrative — Neurath thought protocol sentences were useful, just not foundational — but it isn't the escape from theory-ladenness the piece seems to promise.
Genuine question: does "taste" require something closer to raw causal contact with the lab — Alasdair's living graph, video, sensor logs, the works — rather than any verbalized protocol, six fields or sixty? Or is verbalized, theory-laden judgment actually the only kind of "taste" worth training for, since it's the only kind a scientist can act on either?
Wonderful post. Indeed, the final state is insufficient to understand the intelligence that produced it. Intelligence resides partly in the trajectory through the space of possibilities.
This is strikingly parallel to a problem we have been working on from the enterprise side. I think you are describing trajectories, even if you do not use that term.
The scientific paper is not merely an artefact. It is a lossy compression of the trajectory that produced it. For centuries, that compression was a feature: it made knowledge communicable. But AI changes the economics of memory. The discarded information, the alternatives considered, failed branches, uncertainty, disagreement and belief revision, may now be more valuable to an intelligent agent than the polished artefact itself.
Your six fields are very interesting. Taken together, they begin to describe something more than a better log. They describe the grammar of a trajectory: state → possibility space → decision → rationale → outcome → learning → new state.
We have encountered almost exactly the same problem in organisations. A company leaves behind documents, Slack messages, tickets, commits, presentations and databases. These tell us a great deal about what exists now, but remarkably little about how the organisation arrived there. The rejected proposal disappears. The assumption that proved wrong disappears. The reason option B was selected over A and C disappears. The present contains the survivors of history, but not the alternatives history killed.
This is why I think the browser-history analogy, while excellent, may actually understate the argument. A history tells us where someone went. A trajectory tells us how an agent moved through a possibility space: what it knew, what it could have done, what it expected, what constrained it, what it chose, what happened, and what it learned. The past is valuable because it helps the present become intelligible which enables improved inference about the future.
That distinction between a history and a trajectory becomes especially important for AI.
—If we train models only on successful artefacts, we teach them the destinations without the navigation.
—If we preserve trajectories, we potentially give them access to something much closer to judgement or, as you put it, taste.
There is also a larger infrastructure question here. Perhaps the problem is not merely that we need better training datasets. Perhaps our existing information systems have the wrong primitive for an agentic world. They are extraordinarily good at storing objects and relations: documents, rows, files, entities, graphs. They are much worse at representing organised movement through time and possibility.
Our thesis at White Tree is that trajectory should become a first-class computational object.
That leads to a broader claim: the fundamental data object of the AI era may be shifting from the artefact to the trajectory that produced the artefact.
I’ve linked a short video showing how we are approaching this problem from the enterprise side. The convergence with what you describe here is remarkable. https://www.youtube.com/watch?v=ASU4biYWmuk