Essay

Ground Truth

How a fat ox, a medical algorithm, and a papal conclave taught me to trust the crowd.

Taig Mac Carthy Updated

There is a famous story about a crowd guessing the weight of an animal, and for years I told it wrong. I had it as a pig at an English country fair. The real case is better documented, so let me correct the record.

In 1906, at the West of England Fat Stock and Poultry Exhibition in Plymouth, a fat ox was put on show and fairgoers paid to guess its dressed weight. Francis Galton collected the tickets, ran the numbers, and published the result in Nature the next year as “Vox Populi.” Of roughly 787 usable guesses, the median was 1,207 pounds; the ox’s actual dressed weight was 1,198 pounds. Taken together, the crowd was off by less than one percent.

People call this the wisdom of the crowd, and it is worth being precise about why it works, because it is not magic. Every guess carries its own error, some high, some low. When the guesses are independent (nobody merely echoing one loud voice) and come at the problem from different angles, the high and low errors tend to cancel, and what survives is the signal they shared. Remove the independence and it collapses: a thousand people repeating one rumor are no wiser than the rumor. Aggregation is a method with conditions, not a miracle.

Ground truth is something I build

I did not learn to trust aggregation from a fairground, but from the problem I solve for a living.

I am a medical AI researcher and the Lead Researcher at Legit.Health, and much of my published work circles one stubborn difficulty: before you can train a model to recognize anything, you have to tell it the correct answer. Machine learning calls that the ground truth, and in medicine it is genuinely hard to come by.

My first paper on this, on an automatic urticaria activity score, counts hives on the skin to grade how severe a case is. The obvious way to build the ground truth is to have dermatologists mark every hive in a photograph, and when you do, the specialists disagree. One draws a box where another draws nothing; some annotations are not merely different but mutually exclusive. Yet supervised learning needs one reliable label per image, and anointing a single best expert throws away everyone else and inherits that person’s blind spots.

So we do not anoint one. The method turns each expert’s bounding box into a Gaussian distribution, then combines them in a sum weighted by how well each annotator has performed. What comes out is a lesion-confidence map, a surface highest exactly where the experts most agree, and that map, not any single doctor’s drawing, becomes the label the model learns from. Consensus here is not a mood or a show of hands; it is a specific, weighted aggregation, and it yields a model good enough to be clinically useful.

I used the same instinct again in a later first-author paper on automatically scoring the severity of psoriasis, where the ground truth was again built by aggregating independent expert annotations. The point is that “the crowd, properly aggregated, is smarter than the individual” is not a metaphor I reached for to dress up a music project. It is an engineering method I have implemented more than once, and it works. It is not an appeal to popularity either: weighted by performance and bounded by a method, it is the opposite of counting likes.

Three ways to be less wrong

Once you see aggregation as a method, you notice people have reinvented it wherever the cost of one person being wrong is too high.

The crowd at the ox is the crudest form: no institution, just arithmetic after the fact. Democracy is the same move with structure: rather than let a single ruler’s error become everyone’s fate, many people vote, and in most systems they do not govern directly but elect a representative body, a parliament, to aggregate and deliberate for them. The design assumption is exactly Galton’s, that many independent judgments, pooled, are safer than one.

The version I find most striking is also the strangest, and it is usually described badly. When a pope is chosen, it is not by a vote of all Christians, and it is not direct democracy. A bounded group, the cardinal electors (the cardinals under eighty), is sealed away and votes by secret ballot, and no one wins until a candidate reaches a two-thirds majority. The secrecy protects independence; the supermajority means the result is not a bare fifty-one percent but something nearer real convergence. Whatever you make of the institution, its architecture encodes a claim: the consensus of a disciplined body is a safer route to a right choice than the certainty of any one man.

I am not saying majorities are wise, still less holy. Crowds panic; parliaments pass terrible laws; a conclave is a room of fallible men. None of this proves the many are always right. They are three worked examples of one move: limit the damage a single point of failure can do, by aggregating many independent, varied judgments under a method that weights and bounds them.

Who wrote the Bible

Push this one step further, into the ground the whole project stands on.

Who wrote the Bible? One answer is God. Another is people, thousands of them across thousands of years: mothers telling stories to children by firelight before writing existed, children remembering and retelling and reshaping, stories competing for survival, the ones that lasted because they held truer to something and were harder to forget. Only later were they written, copied, translated, argued over, and sorted by scribes and councils and scholars across centuries.

I cannot make those two answers fight. What if they are one answer seen from two sides, and the thing we point at when we say God is, among other things, the ground truth that surfaces when you aggregate millions of human beings reaching for the same reality across millennia? I do not offer that as proof. I offer it because I have looked hard for the flaw and not found it, and because it makes the Bible more astonishing to me, not less. In the sense that matters most, scripture is a consensus document, the longest-running wisdom-of-the-crowd experiment we have.

But here the conditions from the ox return. Aggregation is only as good as the independence and discipline of its inputs, and this is where I part from the notion that truth comes from pouring in everything. The Bible is not the average of everything anyone ever said; it is a principled selection, material weighed and canonized before it was treated as authoritative, much as my method weights annotators before it trusts them. The music model behind these songs is the same in kind: as I argued in The Machine Is Not the Muse, it was gathered from the open internet, yes, but it is built of created works, not the undifferentiated runoff a language model swallows. That is the difference between this and a model trained indiscriminately on Reddit posts and tweets. The selection is not a detail; it is the point.

The ego dissolved into tradition

I trust the anonymous version of consensus over the authored one, for a personal reason.

I was raised Catholic. Years later I went to a Protestant service and something clicked, though not as intended: I could see the personal taste of whoever had organized it in everything, the decoration, the choice of songs, the small flourishes. A particular human ego, however well meaning, had placed itself between me and the thing I came for, and it made me unexpectedly angry.

Catholic and Orthodox liturgy moves differently. The shape of the space, the order of the rite, the cadence of the chant were settled so long ago, by so many anonymous people, that no one alive can claim them. Nobody knows or cares who chose the angle of the incense; the decisions have dissolved into tradition, and the individual has disappeared. That anonymity felt right to me, and it is close to what I value in music with no author to idolize.

I will be honest about the cost. The same anonymity that dissolves the ego can also calcify. A living consensus hardens into obligation and the life goes out of it; a tradition that only repeats itself becomes a fossil of a conversation that was once alive. Aggregation is a living process or it is nothing.

This leaves the deepest question untouched. Aggregation only works if there is something real for the guesses to converge on. The ox had a true weight; skin has a real number of hives. What is the equivalent for music, or for meaning, and why should there be any truth beneath so much human variation? I cannot claim we fully understand that order; I can only follow the effects it leaves. Those effects are the subject of An Order We Did Not Write, the last of the three essays gathered here.