The brain was in the business of truth and error long before it had words. It still does most of that work without them.
You have done this. You are climbing a staircase in the dark, near the top, and your foot swings up and forward for the next step and finds only floor. The landing arrives early. There is a jolt through the spine, a lurch, half a second of something close to alarm, and then the recognition that you are fine and there were twelve stairs, not thirteen.
Take that moment apart, because a great deal is hiding inside it.
Your leg went to a particular height. Not a vague height: a specific one, matched to the riser of that staircase, held steady across a dozen repetitions in the dark. Your ankle was preloaded to a stiffness suited to the load it expected. Your trunk had already begun tilting to accept a transfer of weight that never arrived. You held, in short, a detailed quantitative commitment about the state of the world. And it was false.
Now notice what did not happen. You never thought there is another step. No sentence formed anywhere. Nothing was believed in the way philosophers usually mean when they say believed. And yet the thing that failed had every property that matters about a belief: it was about the world, it was specific, it was load-bearing, other systems depended on it, and it turned out to be wrong. The lurch in your spine was the error message.
That lurch is the subject of this piece.
The words are the exception
We treat facts as things made of language. A fact is what a true sentence says. Knowledge is justified true belief, and belief is an attitude taken toward a proposition, and a proposition is the sort of thing you could in principle write down. This picture is so ingrained that its alternative can be hard to see even while you are living inside it.
But the picture gets the order of business backwards. Nervous systems were tracking, storing, and acting on the state of the world for something on the order of half a billion years before any of them could say a word about it. A jumping spider stalking prey across a gap it cannot see the far side of is doing something that deserves to be called getting the world right. The spider is committed to facts about geometry it will never utter. So is your ankle.
The claim I want to make is that truth, error, evidence, inference, and memory all have perfectly good non-linguistic implementations, that these implementations are the older and larger part of your own cognition, and that language is a thin and very recent layer sitting on top of a system that was already fully occupied with being right and wrong. When you learn what the older layer is doing, the linguistic case starts to look less like the paradigm of knowledge and more like a specialised extension built for a specific purpose: getting a fact out of one head and into another.
Being right without saying anything
Start with what a fact does rather than what it is made of.
“The floor is solid” earns its keep by making you willing to shift your weight onto your front foot. Strip out the words and the commitment survives intact. Your body has already bet on the floor before any sentence shows up, and it will collect or lose on that bet regardless of whether you ever formulate it.
So the pre-linguistic analogue of a true belief is a well-calibrated expectation, and the pre-linguistic analogue of falsehood is surprise. Not the emotion, necessarily. The measurable quantity: a mismatch between what a system predicted and what arrived.
This is the backbone of predictive processing, an account of cortical function with a long pedigree. Hermann von Helmholtz argued in the 1860s that perception is unconscious inference, that the visual system does not receive the world so much as guess at it, constrained by data. The modern version, developed by Rajesh Rao and Dana Ballard, Karl Friston, Andy Clark, Jakob Hohwy and others, makes guessing the main event. On this view the brain runs a generative model that continuously forecasts its own incoming signals, and the only thing that travels up the hierarchy is the residual, the part the model failed to anticipate. Ie… we notice change. Prediction error is the currency of being wrong. Learning is whatever reduces it.
Critics have pointed out that this kind of reasoning risks explaining everything and therefore forbidding nothing. Treat it as a productive research programme rather than a settled result. But the specific phenomenon underneath it, that brains learn from discrepancy rather than from mere occurrence, is about as solid as behavioural neuroscience gets, and it was established long before anyone drew a predictive coding diagram.
Consider Leon Kamin's blocking effect, from 1968. Teach a rat that a light predicts a shock. Then present the light and a tone together, still followed by shock. The rat learns almost nothing about the tone. The tone was paired with the shock, repeatedly and reliably, and the animal shrugs. Why? Because the light already predicted the shock, so the tone added no information. The animal is not counting co-occurrences. It is tracking whether a cue improves its forecast, which is a primitive and rather good epistemology. Robert Rescorla made the same point from the other direction: what animals learn tracks the genuine contingency between cue and outcome, the difference between how often the outcome appears with the cue and how often it appears without it. Pairings alone do not do it. Correlation is not enough. The animal wants something closer to “information”.
Then, in 1997, Wolfram Schultz, Peter Dayan and Read Montague reported what the discrepancy looks like in spikes. Midbrain dopamine neurons fire in a burst when a reward arrives unexpectedly. Once a cue reliably predicts that reward, the burst migrates backwards in time to the cue, and the reward itself, now expected, produces nothing. And here is the part worth sitting with: if the predicted reward is withheld, those neurons fall silent, dropping below their baseline rate, at the precise moment the reward was due.
The brain knows when to be disappointed. It has an appointment, and it notices the no-show. There is a signal in your midbrain whose meaning is roughly “less than expected, right now,” and it is the closest thing the pre-linguistic brain has to the word not. (It is not really negation, but it is a physically implemented announcement of an absence, which is more than one might have expected to find without syntax.) Later work has complicated this, showing that dopamine carries more than a single error term, but the core finding has held up for nearly three decades.
Facts shaped like maps, facts shaped like reflexes
If truth is calibration, then any format capable of being calibrated can hold facts. Language is one such format. The brain uses several others, and each has its own characteristic way of being wrong.
Maps: Place cells in the hippocampus, discovered by John O'Keefe and Jonathan Dostrovsky in 1971, fire when an animal occupies a particular location. Grid cells in entorhinal cortex, found by Edvard and May-Britt Moser's group in 2005, tile space with a hexagonal lattice that provides something like a metric. Between them they give a rat an enormous body of facts (this is over there, that far in that direction, these two places are adjacent) held not as a list of statements but as a geometry. Edward Tolman had inferred as much in 1948 from behaviour alone: rats allowed to wander an unrewarded maze appeared to learn nothing, but the moment food was introduced their error rates collapsed almost overnight. They had been accumulating structure the entire time, for no reward, with nothing to show for it until it was needed. A map can be wrong by being distorted, and a distorted map produces a confident walk into a wall.
Forward models: Your cerebellum holds a model of your own limb dynamics accurate enough to predict where your hand will be before it gets there, and accurate enough to cancel the sensory consequences of your own movements. This is why you cannot tickle yourself. Sarah-Jayne Blakemore, Daniel Wolpert and Chris Frith demonstrated the mechanism elegantly by putting a short delay between a person's movement and the resulting touch: as the delay grew from zero to about two-tenths of a second, the sensation became progressively more ticklish. The prediction is time-locked to the millisecond, and when reality slips out of register with it, the touch stops being yours.
Affordances: You do not perceive “a horizontal surface at forty-five centimetres.” You perceive sit-on-ability. James Gibson argued that what we see are opportunities for action, facts about the relation between a body and its surroundings, and William Warren later gave the idea a number. Asked whether a step is climbable, people answer accurately, and they answer in units of their own legs: the critical riser height sits at roughly 0.88 of leg length, the same ratio for short people and tall. The staircase you misjudged in the dark was not misjudged in centimetres. It was misjudged in units of “you”.
Magnitudes: Pre-verbal infants and a great many animals represent quantity as a noisy continuous magnitude obeying Weber's law, so that discriminating eight items from sixteen is easy and fifteen from sixteen is not. Six-month-olds pass the first and fail the second. That is numerical fact-holding with no number words anywhere in the building, and it fails in a characteristic way: not by producing the wrong number, but by producing an appropriately fuzzy one.
Episodes: Nicola Clayton and Anthony Dickinson gave scrub jays two things to store for keeping: wax worms, which the birds prefer but which rot, and peanuts, which keep. After a few hours the jays went for the worms. After five days they went straight for the peanuts and left the worms alone. They had retained what they kept in storage, where they put it, and how long ago, and combined the three to draw a conclusion about the present state of a thing they could not see. Researchers call this “episodic-like” memory, and the hedge is deliberate, since nobody can ask a bird whether the remembering feels like anything. What is not in doubt is the informational structure.
The knowledge without an address
Why can you not simply recite any of this?
The answer is architectural, and it is the deepest difference between the two kinds of fact. A sentence is a discrete token. You can retrieve it whole, quote it, hand it to a stranger, hold it at arm's length without endorsing it, embed it inside another sentence, or negate it. It has an address and it has parts.
Non-linguistic facts are stored as adjustments to a process. The knowledge lives in the synaptic weights that shape what a circuit expects and does, distributed across the very machinery that runs. There is no address to look up because storage and processing are not separate things. Asking to “recite your knowledge” is a little like asking a river to hand over its shape.
This is a real, observed effect. After Henry Molaison's medial temporal lobes were partially removed in 1953, he could no longer form new memories that he could speak of. Brenda Milner had him trace a star while watching his hand only in a mirror. Across days his performance improved with the typical learning curve of a healthy person. And each day he approached the task as one he had never seen. The knowledge was going in, accumulating, and improving, with no verbal or “factual” access to it at all.
There is an older and stranger case. In 1911 the Swiss physician Édouard Claparède greeted a patient with severe amnesia while concealing a pin in his palm. When he next offered his hand she refused to take it, and could not say why, eventually offering the observation that sometimes people hide pins in their hands. She had retained the fact. She had no access to it. What she produced instead was a plausible sentence generated after the fact to cover a decision that had already been made somewhere she could not see.
Thinking without sentences
If facts can be held without language, conclusions can be drawn without logic. Two mechanisms do most of the work that inference does in the linguistic case.
The first is that the format performs the inference for free. Kenneth Craik put it like this over 80 years ago: an organism that carries a small-scale model of external reality can try out alternatives inside its head and react to future situations before they arise. The crucial property of a model is resemblance. If your spatial representation shares structure with the space it represents, then shortcuts, detours and transitive relations are simply readable off it. Nothing deduces them. A physical scale model of a bridge tells you true things about the bridge by being shaped like it, not by encoding propositions about it. Representations that mirror the structure of what they represent get whole classes of conclusion thrown in with the format.
The second is simulation. Rats pausing at a maze junction do an interesting thing that Tolman's contemporaries named “vicarious trial and error”: they sweep their heads back and forth between the options, sometimes for several seconds, before committing. In 2007, Adam Johnson and David Redish recorded hippocampal ensembles during exactly those pauses and found the animal's spatial representation sweeping down first one arm and then the other, out ahead of a body that had not moved. During quiet rest and sleep the same circuits replay past trajectories at compressed speed, and sometimes assemble paths the animal has never actually taken.
This is deliberation in rats — comparing options. Options are constructed, run forward, and evaluated, and then one is chosen. There are no words involved.
The dignity of being wrong
Here is the test that separates a real representation from a mere causal link, and it is a test about failure rather than success.
Consider the Müller-Lyer illusion: two lines of identical length, one with arrowheads pointing outward and one inward, one of which looks obviously longer. Now measure them with a ruler, satisfy yourself completely, and look again. Nothing changes. Your visual system continues to assert something your verbal system knows to be false. Two stores of fact, in one skull, are in open disagreement. The perceptual “fact” declines to be corrected by the fact of the words. (Susceptibility does vary across cultures, which suggests the assumptions it encodes are learned rather than innate, but no amount of propositional knowledge dislodges it in an individual.)
Also consider “the frog”. Jerome Lettvin and colleagues showed in 1959 that the frog's retina contains detectors tuned to small dark moving objects, and a frog will readily snap at a BB pellet swung past on a wire. That state is not a different kind of state from the one produced by an actual fly. It is the same state, being false. The frog is not merely unlucky — the frog is mistaken. The frog is wrong without words.
Philosophers including Ruth Millikan and Fred Dretske have argued that this is precisely what makes such states truth-apt in the first place. Something counts as representing the world when other parts of the system have been built, by evolution or by learning, to depend on it as a stand-in for the world. That dependence is what gives it content. Its predictive power of the real world is its “truth”. And the downstream machinery can be misled, because it was designed to lean on a signal that can, in unusual conditions, lie.
A thermostat does not believe anything. A frog's bug detector arguably does, in a thin sense, because a whole apparatus of tongue and timing was shaped by a history in which that signal meant food. And your ankle, on the twelfth step, was wrong in exactly this way. Not broken. Wrong. Again, without words.
What language actually buys you
It's worth being precise about what language contributes, because it is not truth. Truth was already there in the form of calibration.
What language adds is detachability. A fact put as a sentence can be carried away from the situation that produced it. It can be transmitted to someone who was not there, and to someone who will not be born for a century. It can be tagged with a source and a confidence. It can be doubted while being entertained, which is the whole basis of hypothesis. It can be negated, disjoined and quantified. It can be combined with other facts under explicit rules, so that conclusions follow which no single experience delivered.
The format carries messages differently: A picture struggles to say with specificity that “the cat is not on the mat”. Try to represent “either the north or the south route, and if the north, bring rope” in a purely spatial format and the difficulty becomes obvious. Detachability is an enormous gift, and it is the reason a species with fairly ordinary primate cognition ended up with cathedrals and vaccine schedules.
But the layer sits on top. It does not replace what is underneath, and it is not always well informed about it.
Ask a novice cyclist how they turn a corner at speed and they will tell you they lean. What they actually do first is countersteer: they push the handlebar briefly away from the turn, which drops the bike into the lean. Nearly every novice rider does this, and nearly none of them know it. The motor system has the correct physics. The verbal system has a story, and the story does not matter, because the layer that actually turns the bicycle is not consulting it.
Conclusion: the order of events
There is a second epistemology running inside you, and it is the older one. It stores facts as geometries, as timings, as tunings, as readinesses to act. It draws conclusions by running models forward rather than by deducing. It corrects itself continuously, keeping score in the mismatch between what it expected and what arrived. It can be exquisitely right, as when your hand closes on a thrown ball a fraction before the ball is there, and it can be obviously wrong, and it doesn't care what you have been told.
Most of what you know is in this second epistemology. The sentence you could produce about the staircase, if anyone asked, happens aside from the fact. It's a summary composed for communication.
Next time your foot goes down onto a step that is not there, pay attention to the sequence. The lurch comes first. The words come second, reporting on a conclusion you had already reached.