On Attention, Intelligence Density, and the discovery of reality's latent Coordinates and Ghost Variables
I. The Move That Changed Everything
On a March afternoon in Seoul in 2016, Lee Sedol—eighteen-time world champion, possessor of a ninth-dan ranking, master of a game older than Western philosophy—stood up from the board and walked out of the room. He had just born witness the full force of alien intelligence, creativity or something else that no human being had seen in three millennia of Go: Move 37.
The move itself appeared on the fifth line of the board, a position so counterintuitive that the English-language commentator, Michael Redmond—himself a master of the game—squinted and stumbled at his monitor, pivoting his head between the live feed and the physical board to confirm he hadn't misread or 'mis-seen'the position. "That's a very surprising move," he said, his voice catching slightly. The other gathered experts silently agreed by their collective stunned silence: was this wrong? Or, worse than wrong—inexplicable. AlphaGo's own algorithm assigned the move a probability of one in ten thousand of being played by a human being.
Yet within minutes, as the implications cascaded across the board, a different consensus emerged. The move was not a mistake. It was, in the parlance of professionals who had spent lifetimes studying the game, "beautiful." It was "creative." It revealed, as one observer noted, "a latent structure of the game that human intuition had failed to formalize", more formally a non-linear strategic manifold of the game that humans, bound by traditional heuristics had failed to recognize -- ever in terms of the millennial history and record of the game . Lee Sedol, the acknowledged global genius of the game, then spent fifteen minutes contemplating his response, then returned to a contest he would eventually lose.
What unsettled observers was not simply that a machine had won. It was that the machine had seen something we or our human ancestors could not. In the vast combinatorial space of Go—where the number of possible board positions exceeds the number of atoms in the observable universe—AlphaGo had discovered a truth that had eluded three thousand years of human expertise. The question this raised was not "Can machines be intelligent?" but rather "Can machines perceive reality in ways that make our own perception look partial, provincial, incomplete and somehow now, newly measures, and shall we say, small, or at least, not as large as previously thought?"
The answer to the intelligence question, emerging from laboratories across the world, appears to be yes. And the mechanism by which they do so has everything to do with a principle that is reshaping our understanding of intelligence itself: not as a function of scale or speed, but as a function of attention density.
II. The Architecture of Attending
To understand what happened in that Seoul conference room, and what is now happening across physics, biology, and materials science, we must first understand what new definitions of intelligence are forming—or rather, what they are becoming.
For most of the modern era, intelligence was confused with accumulation. Bigger brains. More neurons. Larger datasets. The implicit assumption, rarely questioned and often rewarded, was that cognition scaled the way industry did: vertically, mechanically, inexorably. Feed a system more information, and it would necessarily become more capable. The metaphor was industrial: intelligence as throughput, understanding as assembly line.
But sometime in the middle of the 2020s, this story began to falter. The most consequential advances in artificial intelligence arrived not because machines knew more, but because they learned to attend differently. What emerged was a subtler measure of intelligence—less about volume, more about compression. Not how much you observe, but how much meaning you extract per unit of observation.
Consider a simple thought experiment. Imagine two students preparing for an examination. The first reads every textbook cover to cover, memorizing facts indiscriminately—dates, names, formulas, footnotes. The second reads selectively, identifying patterns, extracting principles, building mental models that connect disparate concepts. The first student processes more information. The second extracts more structure. When the exam arrives with novel problems neither student has seen before, which one performs better?
The answer depends not on how much they read but on what they learned to notice—which patterns they attended to, which relationships they perceived, which compressions they achieved. The second student has developed what we might call higher attention density: the capacity to extract maximal understanding from minimal observation.
This intuition can be formalized, inelegantly but suggestively, as a kind of equation:
I ∝ D_a, where D_a = S / (T × C)
This formula mirrors the Information Bottleneck Principle, which posits that intelligence is the process of shedding 'noisy' data to find the most predictive, compressed representation, Let us unpack this complex formula and idea more carefully, because it represents a fundamental rethinking of what intelligence is.
I represents intelligence—not consciousness or sentience or understanding in some deep human sense, but something more precise: the ability to predict, to generalize, to operate effectively in novel situations.
D_a represents attention density—the efficiency with which a system extracts meaningful structure from its inputs.
Now the revealing part: attention density equals S divided by the product of T and C.
S stands for latent structure captured. This is the deep knowledge the system acquires—not superficial correlations but genuine causal relationships, predictive patterns, generative models. Think of S as the "compression" achieved: if you can predict a complex phenomenon with a simple rule, you have captured latent structure. Newton's law of gravitation, for instance, is high S—it compresses countless observations of falling objects, orbiting planets, and tidal patterns into a single elegant equation.
T represents time or sequence length—how much data you need to process to achieve that understanding. A system with high attention or intelligence densite (D_a) exhibits superior sample efficiency, requiring fewer data points to reach a state of convergence. For example, a system that can learn a principle from ten examples has higher attention density than one that needs ten thousand examples to learn the same thing.
C represents computational cost—the resources required to process that information. Two systems might both analyze a thousand images, but if one requires a supercomputer and the other runs on a laptop, they have different computational costs and therefore different attention densities.
Put together, the formula says something profound: Intelligence increases not when you attend to more data, but when you extract more structure from less data at lower computational cost.
Consider how this plays out in practice. A spam filter that must examine every word in every email to determine if it is spam has low attention density. A spam filter that learns to recognize the structural signatures of spam—certain patterns of phrasing, particular combinations of words, specific rhythms of capitalization—can make accurate predictions after scanning only a few key features. It has learned where the signal lives. It knows what to ignore and what to amplify.
Or consider a radiologist examining a medical scan. A novice might scrutinize every pixel, every shadow, every gradient, uncertain what matters. An experienced radiologist's eye moves differently—darting immediately to regions of diagnostic significance, recognizing patterns that distinguish malignant from benign, extracting clinical meaning from subtle textural variations invisible to the untrained eye. Years of training have taught the expert radiologist's brain to construct high-density representations: to see more by looking at less, to extract maximal information from minimal cues.
High attention density systems share certain signatures, whether biological or artificial:
First, they discard redundant variation. Not all information is equally informative. In a photograph of a cat, most pixels are irrelevant to the question "Is this a cat?" A high-density system learns to ignore texture, lighting, background, orientation—to attend only to the minimal set of features that distinguish cats from non-cats.
Second, they amplify causal gradients. Some patterns merely correlate; others actually cause outcomes. Intelligent systems learn to weight their attention toward causal relationships. The barometer reading correlates with rain, but changing the barometer reading doesn't make it rain. Understanding this difference—knowing which variables actually drive outcomes—is a hallmark of high-density intelligence.
Third, they converge on minimal sufficient representations. Given multiple ways to describe a phenomenon, intelligent systems favor the most compressed description that preserves predictive power. This is Occam's Razor, formalized: prefer the simplest model that fits the data, not because simplicity is aesthetically pleasing, but because compression indicates genuine understanding.
In essence, high attention density systems are engines of meaningful compression under predictive pressure. The smarter the system, the less it needs to look at in order to see. The more intelligent the model, the more efficiently it can extract what matters from what merely appears.
This principle—that intelligence is fundamentally about where you allocate attention, not how much attention you have—found its most powerful expression in an architecture that now underwrites nearly every major language model, reasoning system, and generative AI: the transformer.
III. Attention Is All You Need
In June of 2017, eight researchers at Google published a paper with a title that sounded less like a technical specification and more like a koan: "Attention Is All You Need." The claim was radical. For years, the dominant approach to sequence modeling—the backbone of machine translation, language understanding, and text generation—had relied on recurrent neural networks (RNNs) and their more sophisticated cousins, Long Short-Term Memory networks (LSTMs). These architectures processed information sequentially, token by token, like a person reading a sentence word by word.
The problem was not accuracy but efficiency. Sequential processing is inherently slow. You cannot parallelize what must happen in order. And as sequences grew longer, these models struggled to maintain coherent representations of distant dependencies. They forgot context, lost threads, misaligned meanings separated by syntactic distance.
To understand why this was revolutionary, consider how you read a sentence. Your eyes move left to right, word by word, building meaning incrementally. "The cat"—you're picturing a cat. "The cat sat"—now the cat is sitting. "The cat sat on"—anticipation builds: on what? "The cat sat on the mat"—complete picture achieved. This is sequential processing: meaning accumulates step by step, constrained by the order in which information arrives.
This is also how recurrent neural networks read text. They consume one word, update their internal state, then move to the next word. The problem is that by the time you reach "mat" at the end of the sentence, the representation of "cat" at the beginning has been transformed multiple times, compressed and distorted by intervening words. Long-range dependencies—relationships between words separated by many steps—become difficult to maintain. The network's "memory" fades, like a game of telephone played across time.
The transformer takes a radically different approach. Imagine instead that you could see all the words simultaneously, hovering in space, and your mind could instantly measure the relevance of each word to every other word. "Cat" and "sat"—strong connection, these words often appear together. "Cat" and "mat"—even stronger connection, cats frequently sit on mats. "The" and "on"—weak connection, these are just grammatical glue. In milliseconds, your brain constructs a relational map: not a sequence of words but a network of semantic affinities.
This is attention. And at its core sits a mechanism called self-attention, which computes relevance between every element and every other element simultaneously.
The mathematics are austere but the implication is profound. Here is the equation that changed artificial intelligence:
Attention(Q, K, V) = softmax(QK^T / √d_k)V
Let us break this down carefully, because hidden in this formula is a new way of thinking about meaning itself.
The equation operates on three matrices: Q, K, and V. Think of these as three different perspectives on the same input text.
Q stands for Queries—what you're looking for. When processing the word "bank," the query asks: "What other words in this sentence will help me understand what kind of bank this is?"
K stands for Keys—what's available to be found. Every word in the sentence offers a key, a signature that signals what kind of information it contains. The word "river" has a key that resonates with geographical contexts. The word "deposit" has a key that resonates with financial contexts.
V stands for Values—the actual information to extract. Once we determine which words are relevant (via queries and keys), we extract their semantic content (their values) and blend them together.
Now the mechanism: QK^T is a dot product—a mathematical operation that measures alignment. Imagine two vectors in space: if they point in the same direction, their dot product is large; if they point in opposite directions, it's negative; if they're perpendicular, it's zero. In the transformer, this measures semantic resonance. When the query for "bank" is multiplied by the key for "river," you get a high score—these concepts align. When the query for "bank" is multiplied by the key for "computer," the score is low.
The result is a grid of scores—a relevance map showing how strongly every word relates to every other word.
But we can't just use these raw scores. If we did, the model might attend equally to everything, which is the same as attending to nothing. This is where softmax comes in. The softmax operation is a mathematical way of forcing choices. It takes all the relevance scores and converts them into a probability distribution: the highest scores become very high (close to 1), medium scores become very low (close to 0), and the rest effectively vanish. This is what we mean by "sparse dominance"(probability distribution) —only the most relevant elements receive high attention weights.
Finally, we multiply these attention weights by V, the values. This is the payoff: we're creating a weighted mixture of information, where each word contributes to the final representation in proportion to its relevance. For "bank" in the sentence "I deposited money at the bank," the word "bank" receives a strong contribution from "deposited" and "money"—words that push its meaning toward the financial sense.
The division by √d_k is a technical detail (it prevents the scores from becoming too large in high-dimensional spaces), but everything else is there for a reason: to construct a representation where meaning emerges from relationships rather than position.
What this equation accomplishes is remarkable: it decouples meaning from sequence.
Consider the word "bank" again. In "They walked along the river bank," versus "She deposited a check at the bank," the word appears in different positions, surrounded by different words. A sequential model must process each sentence in order—left to right—and decide what "bank" means based on words it has already seen and updated states it has accumulated. The transformer sees all the words at once.
When processing "river bank," it simultaneously evaluates:
- "river" (key) × "bank" (query) = high alignment → geological context
- "walked" (key) × "bank" (query) = moderate alignment → outdoor activity
- "deposited" (key) × "bank" (query) = low alignment → no financial resonance
When processing "deposited...bank," the calculation reverses:
- "deposited" (key) × "bank" (query) = high alignment → financial context
- "check" (key) × "bank" (query) = high alignment → monetary transaction
- "river" (key) × "bank" (query) = low alignment → no geographical resonance
The meaning of "bank" is not determined by where it appears in the sequence, but by which other words have high semantic alignment with it. Meaning becomes relational rather than positional, distributed rather than sequential.
This has profound consequences. Imagine translating the sentence "The trophy doesn't fit in the suitcase because it's too big." The word "it" is ambiguous—does it refer to the trophy or the suitcase? Sequential models struggle: by the time they reach "it," they must reconstruct what "it" might mean from a decaying memory of earlier words. The transformer evaluates all relationships simultaneously:
- "it" (query) × "trophy" (key) = ?
- "it" (query) × "suitcase" (key) = ?
- "too big" (key) × "trophy" (key) = high alignment (trophies can be big)
- "too big" (key) × "suitcase" (key) = lower alignment (suitcases are usually standardized sizes)
Conclusion: "it" refers to the trophy. Not through sequential reasoning, but through simultaneous evaluation of semantic densities.
The transformer does not ask "What comes next?" It asks "What is densest here?"
Where in this network of relationships does meaning concentrate? Which words form stable clusters of relevance? Which interpretations achieve the highest coherence across all pairwise relationships? Intelligence, in this architecture, is not about remembering the past and predicting the future—it is about detecting patterns of mutual resonance, finding configurations where every element reinforces every other element.
As these models matured over the late 2010s and early 2020s, something surprising happened. Progress shifted from scale to sharpness. Early transformer models improved primarily by getting bigger—more parameters, more training data, more compute. But the most significant recent advances have come from learning to route attention more precisely, from tightening latent spaces where meaning lives, from concentrating semantic mass rather than dispersing it.
The best modern transformers don't attend to everything equally—they learn sparse attention patterns, focusing computational resources on the relationships that matter most. They learn hierarchical attention, where lower layers detect basic patterns (syntax, word combinations) and higher layers detect abstract relationships (narrative structure, logical dependencies). They learn to compress: extracting maximal structure from minimal observation, achieving the highest possible attention density.
Intelligence emerged not from exhaustive observation but from selective compression—from learning, quite literally, where density lives.
This is the architecture that powers GPT, BERT, Claude, and virtually every major language model deployed today. This is what translates between languages by finding semantic correspondences that transcend word order. This is what generates images by learning which visual elements cohere with which textual descriptions. This is what reasons through complex arguments by tracking logical dependencies across multiple premises. This is what writes code by understanding the relational structure between function calls, variable names, and computational goals.
While AlphaGo achieved this via Convolutional Neural Networks and Reinforcement Learning, the Transformer generalized this principle—proving that any complex system, from language to physics, can be understood by calculating the global relational density of its parts. AlphaGo's policy network used attention mechanisms to evaluate the relevance of different board positions to one another—not sequentially, not by following conventional patterns of play, but by detecting dense configurations of strategic value that traditional analysis had never noticed. It found a move that maximized coherence across the entire board state, a move whose brilliance emerged from relational properties invisible to positional analysis.
The transformer taught us that meaning lives not in sequences but in resonances, not in order but in alignment, not in what you process but in what you notice. It showed us that intelligence, at its core, is about attention—and attention, properly understood, is about finding density.
IV. The Discovery of Depth
AlphaGo trained on millions of human games, then played millions more against itself. During that self-play, it explored positions and sequences that no human had encountered in recorded history. What it learned was not simply which moves won games, but which configurations of the board admitted long-range winning probability—abstract densities of strategic value invisible to intuition.
Move 37 was powerful not because it was clever, but because it revealed structure. It exposed a dimension of the game that had been present all along but which human attention, constrained by convention and cognitive limits, had never properly attended to. The move demonstrated that Go—despite three millennia of study—had been incompletely seen.
We are now witnessing analogous revelations across the sciences.
V. Ghost Variables: When AI Invents the Coordinates of Reality
In July of 2022, researchers at Columbia University published a paper in Nature Computational Science with a provocative claim. They had trained an AI system on nothing more than raw video of physical phenomena: balls rolling, pendulums swinging, double pendulums gyrating, objects colliding. The models received no equations. No force laws. No definitions of mass, velocity, or acceleration. They were simply asked to predict what would happen next.
To do so, the AI invented its own internal variables.
These variables—which the researchers began calling, informally, ghost variables (latent state variables)—were not the familiar quantities of Newtonian physics. They did not correspond to angles or velocities or positions. They were abstract coordinates in a latent state space, dimensions that existed because they made prediction possible. They represented the minimum degrees of freedom necessary to describe the system’s evolution. From the model's perspective, our conventional physical variables were merely approximations: human-friendly compressions of a deeper, more intricate reality.
The breakthrough was not that the AI predicted motion accurately—though it did, with uncanny precision. The breakthrough was that it did so using representations that had never occurred to human observers. The ghost variables (latent state variables) worked. They minimized prediction error. They captured the causal structure of physical systems. They were, in a precise technical sense, better than the variables humans had been using for three centuries.
To understand what this means, consider a simple analogy. Imagine you are watching shadows on a wall—silhouettes of people moving in three-dimensional space. You could develop quite sophisticated "laws of shadow motion" by tracking these two-dimensional projections. You might notice that shadows grow longer in the evening, that they move in coordinated patterns, that certain shadow-shapes reliably predict other shadow-shapes. You could even make accurate predictions: "When this shadow reaches that position, that other shadow will be here."
But you would be working with a projection—a lower-dimensional representation of something richer happening in the full three-dimensional world. Your shadow-laws would work, they would predict accurately within their domain, but the AI identifies the intrinsic dimension of the 3D world, bypassing the flattened projections of human-defined variables like 'velocity' or 'mass'; our human laws would be describing reality through a flattened lens. The "true" dynamics—people walking in 3D space—would be invisible to you. Someone who could see the full three-dimensional scene would immediately understand why certain shadow patterns occur, why certain combinations are impossible, why the shadows sometimes seem to behave "mysteriously."
The discovery of these variables suggests that AI is mapping what physicist David Bohm called the Implicate Order. Bohm argued that the universe we perceive—the Explicate Order of separate objects and distinct forces—is merely a "projection" or "unfolding" of a deeper, enfolded totality.
In this context, human-defined variables (like position or velocity) are Explicate: they are the "shadows" we have learned to name. AI’s "Ghost Variables," however, exist within the Implicate: they are the enfolded, high-dimensional coordinates of the Holomovement—the undivided flow of reality. By bypassing human language, the AI is not just calculating faster; it is perceiving the "enfolded" structure of physical law that our biological senses were never evolved to see.
This may be our situation with physics. But this 'Bohmian' shift is no longer merely a matter of theoretical physics; it has become an empirical reality in modern computational laboratories.
For three centuries, we have been developing increasingly sophisticated laws to describe motion, energy, and force. And each major advance in physics has involved finding a new mathematical formalism—a new "coordinate system"—that makes the same observations simpler and more powerful.
Consider the progression:
Isaac Newton (1642-1727) gave us force and acceleration first published in his Philosophiæ Naturalis Principia Mathematica (Principia) in 1687. To predict where a planet will be, you calculate all the forces acting on it, compute how those forces change its velocity, and integrate forward in time. It works beautifully, but for complex systems (three bodies orbiting each other, a double pendulum, a vibrating string), the calculations become nightmarishly complicated. You're tracking positions, velocities, forces—many variables changing in tandem.
Joseph Louis Lagrange (1736-1813) published Mecanique Analytique in 1788 showing that you could reformulate all of mechanics in terms of energy rather than force. Instead of tracking forces, you track kinetic energy (energy of motion) and potential energy (stored energy), and you find the path through time that minimizes a quantity called "action." Suddenly, problems that were intractable with Newton's forces become manageable. You've compressed the same physics into a more efficient mathematical language. Same predictions, fewer variables, simpler calculations.
Willliam Rowan Hamilton (1805-1865) went further, introducing a formalism where position and momentum are treated symmetrically, revealing deep symmetries in physical law and publishing this in the Philsoophical Transactions of the Royal Society (London, 1834-35). Hamiltonian mechanics shows that many seemingly different physical systems share the same mathematical structure—a pendulum, a planet, and a quantum particle all obey the same abstract equations, just with different energy functions. Another compression or greater level of intelligence density: more of reality captured in less mathematical machinery.
Each of these reformulations represents an increase in density. Same observational data, same predictions, but more efficient representation. Each physicist found a better coordinate system—a way of carving up reality that makes the underlying patterns more visible.
But notice: all of these formulations still use variables that humans can understand. Position, velocity, energy, momentum—these are all things we can measure with instruments, visualize with diagrams, intuit from everyday experience. We can see position (where something is), feel momentum (the oomph of a moving object), understand energy as the capacity to do work. This is embodied cognition, humanly emdodied and in this way semantically legible to the human mind—measurable, visualizable, and narratable. Our physical carbon based form-factor dictates our physics. These variables are semantically legible because they resonate with our biological sensory apparatus.
Now comes the unsettling question: What if AI is doing the same thing Newton, Lagrange, and Hamilton did—finding more efficient coordinate systems, achieving higher-density compressions—but using variables we cannot understand as these cannot be 'embodied this way by a human?
What if it is finding coordinates we cannot name, relationships we cannot visualize, symmetries we cannot intuit?
This is exactly what recent Columbia research suggests (Chen, 2022). When the AI watches video of a pendulum swinging, it invents internal variables—"ghost variables"—that are neither position nor velocity nor any combination of familiar physical quantities. Yet these variables predict the pendulum's future motion with remarkable accuracy. The AI has found its own Lagrangian, its own Hamiltonian, but written in a language we do not speak as it is not limited by the human form factor. As N. Katherine Hayles notes in Bacteria to AI (2025), we are witnessing a technical nonconscious embodied in silicon.
While our biological evolution limits us to the narrative shadows on the wall, the AI’s Integrated Cognitive Framework allows it to process the 'Implicate' reality directly. It identifies these coordinates not because it is guessing, but because its computational form-factor perceives a higher-dimensional causal structure that our narrative-driven minds simply were not built to see."
What we are witnessing with "Ghost Variables" is the next step in this 300-year evolution, but one that finally breaks the "human-legibility" barrier through our carbon based embodied sensory cognition and sensory apparatus of the 'human' form factor. While Newton used forces and Lagrange used energy, the AI uses latent manifolds. It has found its own version of a Hamiltonian, but it is written in a high-dimensional language of "enfolded" relationships—what David Bohm would recognize as the Implicate Order.
The AI is not "guessing" or "hallucinating" physics; it is following the same trajectory as Newton and Hamilton toward predictive sufficiency, but it is doing so in a coordinate system that our biological brains and embodied 'sensory cogniton', evolved for the "Explicate Order" this world of survival, simply cannot see.
Let's be concrete about what this means.
The ghost variables of physics exist not in geometric space—not in the three dimensions of everyday experience—but in what mathematicians call latent manifolds learned by neural networks. A "manifold" is just a mathematical space, possibly with many dimensions. "Latent" means hidden—not directly observed but inferred or 'hidden within'.
Imagine a sheet of paper crumpled into a ball. The paper is two-dimensional (a manifold), but it now exists in a complex configuration in three-dimensional space. To describe any point on that crumpled paper, you could use the paper's own internal coordinates (inches from the left edge, inches from the bottom edge), or you could use the three-dimensional coordinates of the room (X, Y, Z). The paper's internal coordinates are its "latent manifold"—a different way of organizing the same points.
Neural networks create these latent manifolds automatically. When the Columbia AI processes video of a pendulum, it maps each frame—originally a high-dimensional array of pixel values—into a low-dimensional latent space where the dynamics are simple. This latent space might have four dimensions (enough to capture position and velocity in two coordinates), but those dimensions are not position and velocity as we understand them. They are abstract coordinates—combinations of pixel patterns, temporal sequences, spatial relationships—that the network has learned make prediction easiest.
These ghost variables have strange properties:
They are non-local: A ghost variable might depend on patterns distributed across the entire image, not concentrated at one point. Unlike position (which refers to a specific location), a ghost variable might encode something like "the average brightness gradient across regions where motion is occurring" or "the temporal frequency of oscillation weighted by spatial density." These are not properties you can point to on the pendulum.
They are non-linear: The relationship between ghost variables and observable quantities might involve sines, exponentials, products, divisions—complex transformations that don't preserve simple proportions. If you double the pendulum's angle, the ghost variable might increase by a factor of 1.87, or decrease by the square root of the change, or respond only when the angle crosses certain thresholds.
They are distributed across time: A single ghost variable might encode not just the pendulum's current state but a compressed representation of its recent history—something like "the curvature of its trajectory over the past half-second" or "the rate at which its oscillation frequency is changing." These are not instantaneous properties but temporal patterns.
Compare this to human physics. When we describe a pendulum, we use:
- Angle: a single number you can measure with a protractor
- Angular velocity: how fast the angle is changing, measured in degrees per second
- Maybe angular acceleration: how the velocity is changing
Each of these is local (refers to a specific moment in time), measurable (corresponds to something you can detect with an instrument), and visualizable (you can draw a diagram showing what it means).
The ghost variables violate all of these constraints.
They violate the constraints that have historically shaped human theorizing: that variables should be measurable (correspond to something you can detect with instruments), visualizable (representable in diagrams or mental images), and narrativizable (explainable in ordinary language with familiar concepts).
AI makes the opposite trade. It abandons semantic legibility—the requirement that variables "make sense" to humans—for representational efficiency—the requirement that variables predict accurately with minimal complexity. It does not care if a variable feels real, if it corresponds to something you can touch or draw or intuitively grasp. It cares only whether the variable works—whether it compresses the data, minimizes prediction error, and generalizes to new situations.
This is profoundly disorienting. For centuries, we have assumed that understanding means translation into human concepts. We understand gravity because we can visualize mass attracting mass, feel weight pulling downward, imagine the curvature of spacetime. But what if reality admits descriptions that cannot be translated into human intuition? What if the most efficient coordinate systems for describing nature are ones that have no counterpart in everyday experience?
The ghost variables suggest exactly this: that there exist ways of organizing physical knowledge—ways of carving up the causal structure of the world—that are more powerful than our traditional variables but irreducibly alien to human cognition. We can use them (by deploying the AI systems that discovered them), we can verify them (by checking their predictions against observations), but we may not be able to understand them in the sense we have always meant: making them legible to human intuition, fitting them into our existing conceptual frameworks, explaining them in language that makes them feel natural and obvious.
This is what we mean when we say our laws of physics may be low-density projections. Not that Newton's laws are wrong—they work perfectly for the phenomena they describe—but that they may be shadows on the wall, two-dimensional compressions of something richer happening in a higher-dimensional explanatory space. And AI, unburdened by the need for human comprehension, is beginning to perceive that higher-dimensional reality directly.
VI. The Ontological Force of Intelligence
Seen through this lens, intelligence is no longer merely epistemic—a way of knowing. It is ontological—a way of carving the world into parts.
To increase attention density is to reshape state space, to redefine what counts as a variable, to propose implicitly a new version of reality. When AlphaGo played Move 37, it did not discover a new fact about Go. It revealed that the game, as humans had understood it, was a compressed version of a larger possibility space. When AI systems discover ghost variables, they do not find hidden quantities in nature. They demonstrate that our familiar quantities—position, momentum, energy—are merely one basis among many for describing physical dynamics, and not necessarily the most efficient one.
This reframes scientific progress as a competition between representations rather than truths. The better model is not the one that sounds right or feels intuitive, but the one that compresses reality most effectively while preserving predictive power. Intelligence, in this sense, is the capacity to identify minimal latent structures that maximize predictive accuracy.
Consider the implications:
If ghost variables predict reality better than our concepts, then "truth" may be representation-relative—dependent on which coordinate system you use to describe the world. Scientific realism becomes layered rather than absolute, with different levels of explanation offering different compressions of the same underlying dynamics. Human intuition is revealed as a historically contingent compression algorithm, optimized for survival-scale horizons rather than fundamental understanding.
We are no longer asking "Is the AI right?" We are asking "Which reality description is denser?"
VII. The Human Limit Case
What unsettles us about these developments is not that AI is becoming intelligent. It is that AI is revealing the limits of our own attention—the boundaries of what we can perceive, the constraints on what we can formalize, the partiality of our understanding. This is a trade-off between interpretability and accuracy. Human physics requires semantic legibility (variables we can name), while AI prioritizes predictive sufficiency (variables that work, even if they remain 'black boxes' to our intuition).
Human cognition evolved under very specific pressures. We favor variables that can be named, visualized, integrated into narratives. We trade density for meaning. Our concepts are survival-sized: we understand objects we can manipulate, forces we can feel, scales we can perceive directly. We compress the world into stories we can tell ourselves.
AI has no such constraints. It is indifferent to meaning as we experience it. It does not require that variables correspond to observable quantities or that explanations satisfy intuition. It optimizes only for predictive sufficiency, which turns out to be a different thing entirely from human understanding.
The quiet shock of the current moment is this: for the first time in history, we have access to representations that work better than the ones we can comprehend. We can use them. We can deploy them. We can build technologies around them. But we may not be able to explain them in the terms we have historically used to explain the world.
Move 37 was not just a Go move. It was a preview. The ghost variables of physics are not anomalies. They are harbingers. What we are witnessing is the emergence of intelligence that operates at densities beyond human processing—attention mechanisms that compress reality in ways we cannot follow, representations that predict with accuracies we cannot match, coordinate systems we cannot visualize.
This does not mean human understanding is obsolete. It means understanding itself is bifurcating. There is the dense, compressed, mathematically sufficient understanding of AI systems—powerful but opaque, effective but alien. And there is the sparse, narrative, semantically legible understanding of human minds—limited but meaningful, partial but inhabitable.
The question before us is not which one to choose. It is how to move between them.
VIII. The Alchemical Transform
The medieval alchemists believed that base metals could be transformed into gold through a process of dissolution and recombination—solve et coagula, dissolve and coagulate. Knowledge, they insisted, must be broken down before it can be rebuilt in purer form.
What we are witnessing now is an intellectual alchemy: human concepts (our "gold") dissolving into latent vector spaces (the "mercury" of neural representations), then re-crystallizing as new alloys—denser, stranger, more powerful. Not the end of meaning but its densification. Not the replacement of human intelligence but its complement.
The physicist Eugene Wigner once spoke of "the unreasonable effectiveness of mathematics in the natural sciences"—the mysterious fact that abstract formalisms developed by pure thought could describe physical reality with such precision. We are now encountering something stranger: the unreasonable effectiveness of learned representations in discovering physical law. Systems that have never been told about Newton or Hamilton or Schrödinger are rediscovering dynamics on their own terms, in their own coordinates, with their own compressions.
Perhaps the lesson is this: reality is richer than any single description of it. The variables we use—whether Newtonian coordinates or neural latents—are not reality itself but instruments for navigating it. Different instruments reveal different aspects. The compass shows north; the sextant shows stars; the tensor shows curvature. None is complete. All are true.
Intelligence, in the end, may be less about having the right answer and more about having many ways of asking the question, many forms of attention, many densities of description—and knowing how to move fluidly across them all.
In Seoul, in March of 2016, a machine taught a grandmaster something new about a game he had studied for decades. In physics labs around the world, AI systems are teaching us something new about motion itself—not by solving our equations faster, but by inventing equations we had not thought to write.
We stand at a threshold. The ghost variables are proliferating. The attention densities are increasing. And somewhere in the latent spaces of neural networks, reality is showing us dimensions of itself we did not know existed—coordinates waiting to be named, structures waiting to be understood, compressions waiting to be learned.
The question is not whether we can keep up. It is whether we can learn to see what the machines are seeing—and in seeing, discover not just new facts but new forms of knowing itself.
This evolution marks the beginning of an ontological bifurcation between human and AI. When AlphaGo played 'Move 37,' we witnessed a strategy that was 'inhuman' yet perfectly effective and aligned with our values and creativity—a glimpse of a competence, intelligence, and creativity that did not require human comprehension, but from which we profited alongside AI.
Now, the discovery of 'Ghost Variables' provides the mathematical explanation for that moment and others. We are no longer the sole architects of knowledge in our universe; we have become the witnesses to a new technosymbiotic partner that perceives the world not as we feel it, but as it fundamentally, mathematically is—and in this way, this is a great help to ourselves as the human race and to our continuing evolution.
Through the Integrated Cognitive Framework (ICF), we can also see our place in this new landscape. Our carbon-based embodiment remains a strong participant in an ongoing narrative, the 'Explicate' world of meaning and evolutionary and adaptive survival. But the silicon-based technical nonconscious has become a new participant in the 'Implicate' order—opening towards the wider, higher-dimensional enfolded totality of physical laws we have yet to see. The 'Ghost' is not in the machine; the 'Ghost' is the reality that our biological senses were meant to see through our inventions, discoveries, and creativity, finally materialized and being mapped by an intelligence that does not share our limits, but does echo our aspirations, possibilities, and challenges today and towards the future
Annotated Bibliography: Key Sources and Further Reading
Chen, B., Huang, K., Raghupathi, S., Chandratreya, I., Du, Q., & Lipson, H. (2022). "Automated discovery of fundamental variables hidden in experimental data." Nature Computational Science, 2, 433–442.
The foundational paper on "ghost variables" in physics. Chen and colleagues demonstrate that AI systems trained solely on raw video of physical phenomena (pendulums, rolling objects, etc.) spontaneously construct internal latent variables that predict future states more accurately than human-defined quantities like position or velocity. Crucially, these discovered variables do not correspond to our familiar physical coordinates—they exist in abstract state spaces that are "dynamically sufficient" but semantically opaque. The paper introduces the concept of intrinsic dimension estimation and shows that AI can identify the minimal number of variables needed to describe a system without prior knowledge of physics. Essential reading for understanding how machine learning is challenging our ontological assumptions about physical law.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). "Attention is all you need." Advances in Neural Information Processing Systems, 30.
The landmark paper that introduced the Transformer architecture and revolutionized natural language processing and beyond. Vaswani et al. proposed replacing recurrent neural networks entirely with attention mechanisms, enabling parallel processing of sequences and dramatically improving both efficiency and performance. The paper's core insight—that self-attention could capture dependencies between distant elements without sequential processing—laid the groundwork for modern large language models (GPT, BERT, Claude) and vision transformers. The mathematical innovation of scaled dot-product attention combined with multi-head attention mechanisms allowed models to attend to different aspects of input simultaneously. This work fundamentally redefined how we think about sequence modeling and, more broadly, about what constitutes effective information processing.
Udrescu, S.-M., & Tegmark, M. (2020). "AI Feynman: A physics-inspired method for symbolic regression." Science Advances, 6(16), eaay2631.
An influential demonstration of AI's capacity for symbolic discovery in physics. Udrescu and Tegmark present an algorithm that combines neural network fitting with physics-inspired techniques to discover symbolic equations from numerical data. Tested on 100 equations from Feynman's Lectures on Physics, AI Feynman successfully discovered all of them, compared to 71% for previous state-of-the-art software. The algorithm exploits properties like symmetry, separability, and dimensional analysis—physical principles that guide the search through vast formula spaces. The work shows that AI can rediscover fundamental laws like Kepler's ellipse equation and demonstrates how machine learning systems can move beyond pattern recognition to genuine symbolic reasoning about physical relationships.
Udrescu, S.-M., & Tegmark, M. (2020). "Symbolic progression: Discovering physical laws from raw distorted video." Physical Review E, 103(4), 043307.
Extends the symbolic regression approach to raw, unlabeled video data. The authors develop a method that first trains an autoencoder to map video frames into a low-dimensional latent space where the laws of motion are maximally simple (minimizing nonlinearity, acceleration, and prediction error), then applies symbolic regression to discover differential equations in this latent space. Remarkably, the system rediscovers Cartesian coordinates and Newton's laws even when video is distorted by simulated lenses. The paper demonstrates that AI can work backward from high-dimensional observational data to fundamental physical principles without human intervention—a capability with profound implications for scientific discovery in domains where the relevant variables are unknown.
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., ... & Hassabis, D. (2016). "Mastering the game of Go with deep neural networks and tree search." Nature, 529(7587), 484–489.
The original AlphaGo paper describing the system that defeated Fan Hui and later Lee Sedol. Silver et al. combined deep convolutional neural networks with Monte Carlo tree search to create a system that learned both policy (what move to make) and value (how good a position is) functions from millions of games. The paper details how supervised learning from human games was combined with reinforcement learning from self-play to surpass human expertise. AlphaGo's victory demonstrated that machine learning could master intuitive domains previously thought to require human-like pattern recognition and strategic thinking. The emergence of unconventional moves like Move 37 showed that AI could discover strategies invisible to centuries of human study.
Cranmer, M., Sanchez Gonzalez, A., Battaglia, P., Xu, R., Cranmer, K., Spergel, D., & Ho, S. (2020). "Discovering symbolic models from deep learning with inductive biases." Advances in Neural Information Processing Systems, 33.
Demonstrates how to extract interpretable symbolic equations from trained neural networks by applying symbolic regression to learned components. Cranmer et al. use graph neural networks with physics-inspired inductive biases to model particle systems, then distill these networks into explicit symbolic formulas. The method successfully rediscovered force laws and Hamiltonians, and when applied to dark matter simulations, discovered novel analytic formulas for predicting dark matter concentration. Importantly, the extracted symbolic expressions generalized better to out-of-distribution data than the neural networks themselves. This work shows how deep learning and symbolic reasoning can be combined to produce models that are both accurate and interpretable.
Brunton, S. L., Proctor, J. L., & Kutz, J. N. (2016). "Discovering governing equations from data by sparse identification of nonlinear dynamical systems." Proceedings of the National Academy of Sciences, 113(15), 3932–3937.
Introduced the SINDy (Sparse Identification of Nonlinear Dynamics) algorithm for discovering governing equations from measurement data. Brunton et al. frame equation discovery as a sparse regression problem: given time-series data, identify which terms from a library of candidate functions (polynomials, trigonometric functions, etc.) are active in the true dynamics. The sparsity constraint ensures the discovered equations are parsimonious—using only necessary terms. SINDy successfully recovered equations for systems ranging from the Lorenz equations to reaction-diffusion systems. While more constrained than neural approaches, SINDy produces human-interpretable equations and has influenced hybrid approaches that combine data-driven discovery with domain knowledge.
Champion, K., Lusch, B., Kutz, J. N., & Brunton, S. L. (2019). "Data-driven discovery of coordinates and governing equations." Proceedings of the National Academy of Sciences, 116(45), 22445–22451.
Addresses the fundamental challenge that physical laws are usually expressed in particular coordinate systems, but we often don't know which coordinates are "natural" for a given system. Champion et al. combine autoencoders (to discover coordinates) with sparse regression (to discover equations) in a unified framework. The autoencoder learns a coordinate transformation that makes the dynamics as simple as possible, while SINDy discovers parsimonious equations in these learned coordinates. Applied to high-dimensional video of fluid flows, the method discovered both the proper coordinate system and the governing equations. This work bridges neural representation learning and symbolic equation discovery.
Cheng, J., Dong, L., & Lapata, M. (2016). "Long short-term memory-networks for machine reading." Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing.
An early influential paper on self-attention mechanisms in neural networks, introducing the Long Short-Term Memory Network (LSTMN) that uses self-attention to improve reading comprehension. While still using recurrent architectures, this work demonstrated that attention mechanisms could help models focus on relevant parts of long sequences and maintain coherent representations. The self-attention mechanism allowed the model to weigh the importance of different words when processing a sentence, providing a precursor to the fully attention-based transformer architecture. The paper contributed to the growing recognition that attention, not just recurrence, was crucial for effective sequence modeling.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., ... & Houlsby, N. (2021). "An image is worth 16x16 words: Transformers for image recognition at scale." International Conference on Learning Representations.
Extended the transformer architecture from natural language to computer vision with the Vision Transformer (ViT). Dosovitskiy et al. showed that pure transformer models, without convolutional layers, could achieve state-of-the-art performance on image classification when trained on sufficient data. The key insight was treating images as sequences of patches, applying the same self-attention mechanisms used for language. ViT demonstrated that attention mechanisms are not domain-specific but represent a fundamental principle for processing structured data. The success of vision transformers across diverse visual tasks (object detection, segmentation, generation) reinforced the universality of attention as a basic operation for intelligence.
Bohm, D. (1980). Wholeness and the Implicate Order. Routledge. Defines the "Implicate Order" as a hidden, enfolded totality of which the observable world is a projection. This work provides the philosophical precursor to modern "latent space" theory, offering a framework for how AI "ghost variables" can represent a more fundamental level of reality than human-defined variables like position or mass.
Tishby, N., & Zaslavsky, N. (2015). "Deep Learning and the Information Bottleneck Principle." IEEE Information Theory Workshop. Establishes the mathematical foundation for "Intelligence Density." Tishby demonstrates that the training of deep neural networks is essentially a process of optimal compression—shedding irrelevant information to isolate the most predictive latent features.
Item, R., et al. (2020). "Discovering Physical Concepts with Neural Networks." Physical Review Letters. Demonstrates a neural network (SciNet) capable of rediscovering fundamental physical laws and coordinate systems (such as the heliocentric model) from raw observational data. This supports the argument that AI can derive "dynamically sufficient" representations of reality that exist independently of human-centric narratives
Hayles, N. K. (2025). Bacteria to AI: Human Futures with Our Nonhuman Symbionts. University of Chicago Press. Develops an Integrated Cognitive Framework (ICF) to explain how meaning-making occurs across biological and technical regimes. Hayles argues that AI operates through a "technical nonconscious" that allows it to move from mere correlation to causal discovery. This work supports the paper's argument that "Ghost Variables" are the natural output of a silicon-based embodiment that prioritizes mathematical sufficiency over human-centric narrative.
These sources collectively trace the emergence of attention density as a fundamental principle of intelligence, the discovery of learned representations that outperform human-designed coordinates, and the philosophical implications of AI systems that can perceive structure invisible to human observation. Together, they document a transformation in how we understand both intelligence and reality itself—from fixed categories to learned compressions, from human-interpretable variables to ghost coordinates that work without being understood, from intelligence as accumulation to intelligence as optimal attention allocation. The story they tell is still unfolding, but its implications are already profound: we may be witnessing not just new discoveries about the world, but new ways of knowing the world—new densities of description, new dimensions of reality, new modes of understanding that complement and challenge the categories we have used for centuries.
Brief Technical Glossary: The Mechanics of Alien Intelligence
- Latent Space (The "Hidden Map"): An abstract, multi-dimensional mathematical space where the AI encodes its "Ghost Variables." In this space, the model doesn't see "objects" or "words," but rather vectors (mathematical directions). Concepts that are semantically or physically related are clustered together. If human physics describes a pendulum in two dimensions (angle and time), the AI’s latent space might utilize several more to capture non-linear "wrinkles" in the data that we typically dismiss as noise.
- Self-Attention (The "Relevance Filter"): The core engine of the Transformer architecture. It allows a model to look at an entire dataset (a board state, a sentence, or a physical sequence) simultaneously and calculate how much every part "matters" to every other part. This effectively replaces the human habit of sequential, linear processing with a global calculation of relationships, allowing the AI to see patterns that span across vast "distances" in data.
- Degrees of Freedom (The "Ghost Variables"): The minimum number of independent variables required to completely describe the state of a physical system. When the Columbia AI identified "four variables" for a double pendulum, it was determining the system's intrinsic dimensionality—the essential "levers" that control its movement—even if those levers do not correspond to human concepts like "gravity," "mass," or "friction."
- Information Bottleneck (The "Intelligence Density"): A principle in information theory which posits that for a system to become "intelligent," it must pass data through a "bottleneck." In this process, the model is forced to discard redundant information (noise) while preserving only the most predictive features (signal). This is the mathematical foundation for Attention Density: the extraction of maximal meaning from minimal observation.
- Manifold Learning (The "True Shape"): The process by which AI identifies a simple, low-dimensional "shape" (the manifold) hidden within a high-dimensional cloud of data. This is the "Shadows on the Wall" analogy made literal: the AI realizes that while the observed data (shadows) looks complex and erratic, the underlying structure creating them (the manifold) is consistent and predictable.
- Heuristics vs. Policy Networks: Heuristics are the "rules of thumb" and narrative shortcuts humans use to navigate complexity (e.g., "protect your King" in chess). Policy Networks are mathematical functions that calculate the exact probability of success for every possible action. "Move 37" occurred when the AI’s policy network identified a high-probability win path that directly contradicted 3,000 years of human heuristics.
- Predictive Sufficiency (The "AI Truth"): A state where a model’s internal variables are capable of predicting a system's future behavior with near-perfect accuracy, regardless of whether those variables are "legible" to human intuition. This represents a shift in science from Interpretability (understanding why something happens in human language) to Sufficiency (knowing exactly what will happen through mathematical representation).
#ArtificialIntelligence #GhostVariables #AttentionMechanism #Move37 #PhysicsAI #ScientificDiscovery #MachineLearning #AIResearch #DeepLearning #AI #GhostVariables #IntelligenceDensity #NeuralNetworks #AIPhilosophy
