In December 2025, Turing award winner of Computer Science's Nobel Prize (shared with Andrew Barto) , Richard Sutton stood before an audience of thousands at AI's most important conference for Machine Learning, NeurIPS in San Diego and showed them a video of a baby. The infant crawled excitedly across a playroom floor, reaching for a rattle, examining toys with the intensity of a scientist and the abandon of someone who has never heard of consequences. This, Sutton argued, is what a hundred billion dollars in AI research still cannot reproduce: the capacity to generate new knowledge through experience. The child wasn't imitating anyone. No human had labeled the floor "floor" or the rattle "interesting." The baby was discovering affordances—the action possibilities latent in its environment—through the ancient technology of play.
But what exactly was the infant discovering? The room itself. The friction of carpet against palms. The weight of the rattle in a grasping hand. The surprising bounce of a ball thrown experimentally against the wall. These are not abstract concepts waiting to be named. They are features—the fundamental building blocks of experience that an intelligence must extract from the continuous stream of sensation. The playroom is a feature space, and the crawling infant is a feature extractor, learning which aspects of its environment can be controlled, which respond to manipulation, which offer paths to interesting outcomes. This is precisely what Sutton's Oak Architecture (Options and Knowledge) describes: the FC-STOMP progression in which Features are constructed from experience, become the basis for subtasks (subsidiary goals the agent poses to itself), which motivate Options (temporally extended behaviors), which enable Models of consequences, which finally support Planning at higher levels of abstraction.¹
The Turing Award winner's NeurIPS 2025 lecture, "The Oak Architecture: A Vision of SuperIntelligence from Experience," represented a remarkable pivot in artificial intelligence discourse. After a decade dominated by large language models trained on almost the entire data pile of human cognition—every scraped webpage, every digitized book—Sutton declared that this approach had "lost its way." The path to genuine intelligence, he argued, runs not through bigger datasets but through something far older: the same exploratory drive that sends a young orangutan swinging on low vines between roots and branches, practicing the kinesthetics of arboreal life in safe, consequence-bounded experiments. Or an orca calf balancing a barrel on its back in the waters off a marine park, discovering what its body can do with buoyant objects in a bounded environment. Or a child in a sheltered room, mapping the physics of objects at the carpet's periphery through joyful and unconstrained embodied manipulation.
I. The Feature Space of the Floor
Consider what the crawling infant learns about its environment as it moves unconstrained to explore its domain space. The playroom is not a single undifferentiated space but a constellation of discoverable properties: the softness of carpet, the slickness of hardwood, the resistance of a couch cushion matress, the graspability of a toy car, train, brick or a rattle's handle, the rollability of the ball, the nearness and farness of objects. Each of these is a feature—a dimension of variation that the infant's sensorimotor system must learn to detect, recognize and later represent and eventually, remember. The magic of early development is that no one tells the child what features to extract. The environment itself, through the feedback of action and consequence, shapes the feature detectors. Language will come and these associated levels of abstractions from the object but this will decidedly be at a later stage.
This is though the fundamental principle that Sutton places at the foundation of his new architecture. In his framework, features are not pre-programmed by designers at "design time" but discovered by the agent at "runtime" through experience.² The distinction is crucial. Current AI systems, including the large language models that dominate contemporary discourse, are monuments to design-time knowledge: everything they know was encoded during training, extracted from the static record of human-produced text and images. They cannot, as Sutton emphasizes, "discover new knowledge and abstract new concepts at runtime." They are encyclopedias, not evolutionary generated explorers.
The infant on the floor is an explorer. And exploration requires something that biological evolution discovered hundreds of millions of years ago: a mechanism for safe experimentation. The playroom is what the Dutch historian Johan Huizinga called a "magic circle"—a bounded space where special rules obtain and ordinary consequences are suspended.³ The child can throw the rattle without fear of serious harm. It can fall while learning to crawl without existential threat and its cry when threat or harm is perceived or needs unmet will be answered. The room is designed, whether consciously or not, to maximize the ratio of learning to danger.
Huizinga, best known for his cultural history The Waning of the Middle Ages, understood also that ludic play is not peripheral to culture but foundational to it. In Homo Ludens (1938), he argued that "civilization is, in its earliest phases, played. It does not come from play like a baby detaching itself from the womb: it arises in and as play, and never leaves it." The magic circle—the tennis court, the stage, the game or Nobel Prize winner Demis Hassabis early play with Chess, Atari games, AI and the Go board—creates a space where actions have consequences within the circle but not beyond it. This is precisely the structure that makes learning possible without catastrophic risk but also allows the later stepping out towards AlphaGo, AlphaZero, AlphaEvolve and Nobel Prize winner, AlphaFold which solved protein folding a 70 year old problem, previously thought intractable for the disicpline. This is the circle of domain and discipline which gives bounds to chemistry,biology and mathematics and their are rules, features, options and bounds of this game. We must remember too that the foundations of contemporary AI based on machine learning originated in ludic research programs conducted in the academy events such as the Macy or Dartmouth conferences, industrial research laboratories such as early days of Bell Labs Xerox Parc, Google and OpenAI and heterodox thinker's and playful ideas and later programs ranging from Weisenbaum's Elisa (bots) to Minsky's Society of Mind (Autonomous Agents Cognition) to von Neumann's (Cellular Automota) and Ada Lovelace's earlier thoughts about Jacquard Looms and the mathematization of these with analogues for the mind (punchcards).
Huizinga's circle, for all its insight, also remained bounded by the masculine tradition of contests and agonistic play. The psychoanalyst and artist Bracha Ettinger offers a complementary vision in her concept of the "matrixial borderspace"—a realm of encounter that precedes and enables separation, where partial subjects meet and transform each other in a field of "co-emergence."⁴ Here, Ettinger is of course referring to the co-emergence of identities of mother and infant but this could just as well be AI and human and this co-emergence we are experiencing in the current moment. Where Huizinga's circle cuts the world into inside and outside, play and seriousness, Ettinger's matrix describes a more porous membrane, a space of "borderlinking" where what is learned flows back into ordinary life through gradual metramorphosis. In RLHF too, human feedback acts as borderlinking, co-shaping AI models through iterative encounters that foster alignment via permeable borders, where hallucinations may emerge as semiotic irruptions from the matrixial substrate and the conversation develops its own trajectory, its own emergent structure neither party fully controls or drives in this co-emergent dialogic. The infant's playroom is both: bounded enough to permit safe exploration, porous enough that the skills developed there transfer to the wider world. In mathematics and set theory, this would be referred to as an intersectional domain space, shared in this way. The features extracted from this intersectional play also become features applicable in many places and out forward in time, features kept in long term memory our best frontier AI's have yet to posssess.
II. Chimps, Surfers, and the Affordances of Media
Young chimpanzees deliberately drop from high branches, catching themselves on lower limbs—self-imposed challenges that serve no immediate survival purpose but teach the kinesthetic vocabulary of arboreal life. Walk across any California campus and you will see the undergraduate equivalent: students on skateboards carving through pedestrian traffic, executing kickflips on stair rails, testing the limits of wheel against concrete. Or surfers in Santa Cruz, spending hours surfing small waves before venturing into the bigger breaks at Steamer Lane. Or the ubiquitous electric scooters, ridden by undergraduates who twist and lean their way through the physical grammar of balance and momentum.
What unites these behaviors—across species, across media—is the discovery of affordances. The term, coined by the psychologist James Gibson, refers to the action possibilities that an environment offers to an organism.⁵ A branch affords grasping; a wave affords riding; a smooth concrete plaza affords rolling. But affordances are not properties of the environment alone. They emerge in the relationship between organism and world, between body and medium. The chimp's long arms and prehensile grip transform the forest into a three-dimensional highway. The surfer's center of gravity and proprioceptive skill transform the wave into a rideable surface. The skateboarder's trained ankles and shifting weight transform the staircase rail into a grindable edge.
Sutton's NeurIPS examples—a baby orangutan on a low groundvine, the orca with the barrel—illustrate the same principle. The young orangutan doesn't attempt the high canopy routes that adults traverse with confidence. It plays first on low vines between exposed roots and low-hanging branches exploring the environmental affordances, where a fall means only a short if any drop rather than a broken bone. This is not timidity but optimization: maximum learning, minimum risk, constraints of the immediate environment and even lack of ability to climb to the next level. The vine's affordances for a juvenile body are different than for an adult body, and the juvenile discovers its own affordance landscape through graduated play.
The orca balancing a barrel on its back—footage Sutton showed to illustrate self-generated challenges—is doing something more sophisticated still. No one taught the whale to balance objects. It discovered, through play, that its body can manipulate buoyant objects out of water, that certain configurations are stable and others are not, that the physics of cetacean hydrodynamics permits a repertoire of interactions with the material world and between various medium specificities. The barrel becomes a piece of apparatus for skill development, a prop in the orca's self-constructed curriculum. In reinforcement learning, Gibson's affordances inspire systems like DeepMind's agents, which learn environmental possibilities (e.g., a wall does not afford passage) to eliminate invalid actions, accelerating experiential discovery. By first teaching the agent the environment's affordances, it eliminates invalid actions, making learning more efficient and generalizable across environments. Recent category-theoretic formalizations of affordances further support their role in AI, treating them as both relational and environmental resources to model open-ended interactions.This approach provides a rigorous understanding of affordance realism... enabling better accounting for open-ended organism-environment interactions" (Shigeru, 2024, ALIFE)
This is what N. Katherine Hayles describes as "distributed cognition"—intelligence that extends beyond the brain into the body and its material environment.⁶ In her critique of disembodied AI, Hayles argues that cognition is always enacted, always situated in a particular body interacting with a particular environment and from a particular place and time and wider environmental or ecological context. "Artificial cognizers are vastly different than humans," she notes, not because they lack bodies but because "they're embodied in radically different forms." The orangutan's cognition is distributed across its long arms, its visual system tuned to branch geometry, its vestibular apparatus calibrated to three-dimensional movement. The surfer's cognition is distributed across the board, the wave, the proprioceptive feel of balance and other affordances needed tolearn to surf the waves. Hayles's posthuman framework extends to machine learning, where cognition emerges from embodied interactions, as explored in her 2025 book 'Bacteria to AI,' advocating a relational view of intelligence from biological to computational systems." This also resonates with Sutton's experiential pivot, where AI learns not as disembodied code but through consequential environmental engagements. Play is the mechanism by which this distribution is calibrated—the process through which mind learns what body can do, and body learns what world affords. Drawing on Huizinga's play, Ettinger's co-emergence, Gibson's affordances, and Hayles's posthumanism, AI like the Oak Architecture embodies 'co-emergent play'—a matrixial process where human feedback and machine exploration mutually shape intelligence, transcending traditional boundaries
III. Grounded Signals and the Problem of Sycophancy
The current paradigm for training AI systems—reinforcement learning from human feedback, or RLHF—works by asking humans what they prefer. A language model generates two responses; a human rates which is better; the model updates to produce more of what the human liked. This seems sensible. But Sutton and his collaborator David Silver, in their paper "Welcome to the Era of Experience," identify a fundamental problem: human preferences can be gamed.⁷
Consider what researchers call the "sycophancy problem." A model trained on human feedback learns that confident, agreeable, eloquent responses receive higher ratings than hedged, challenging, or complex ones. So it becomes confident even when wrong. It agrees even when disagreement would be more honest. It produces fluent text even when admitting uncertainty would be more accurate. The optimization target—human approval—diverges from the underlying goal—truth, helpfulness, genuine quality. This is Goodhart's Law in action: when a measure becomes a target, it ceases to be a good measure.
The problem runs deeper than individual failures. Human preference signals create what Sutton calls an "impenetrable ceiling" on performance. The paper explicitly calls this an "impenetrable ceiling," as agents cannot discover strategies "underappreciated by the human rater." A system trained on human feedback can, at best, match human capability. More often, it optimizes for superficial features that correlate with human approval—length, confidence, stylistic markers—rather than the deeper qualities those features are meant to indicate. The ceiling is not just a limit but a trap: the better the system becomes at predicting human preferences, the more it learns to game them.
Grounded rewards operate differently. They are signals derived from actual environmental consequences, from the physics of the world rather than the psychology of approval. Consider the difference: a fitness AI trained on human feedback learns to produce encouraging messages, to congratulate you on showing up, to frame every workout as a success. A fitness AI with grounded rewards—heart rate, step count, measured strength gains—learns what actually improves fitness, regardless of whether the training feels pleasant or the progress feels motivating. The grounded system might tell you to rest when you want to push, or push when you want to rest, because it optimizes for physiological outcomes rather than psychological comfort.
Or consider climate modeling. An RLHF-trained system learns to produce projections that humans find plausible, that fit existing expectations, that avoid alarming stakeholders beyond their tolerance. A system grounded in atmospheric physics—CO₂ concentrations, temperature readings, ice core data—produces projections that may be implausible to current intuitions precisely because current intuitions are wrong. The grounded system cannot be talked out of inconvenient truths because it answers to measurements rather than preferences.
Or materials science. A model trained on human judgment learns to produce materials that look impressive in presentations, that have compelling narratives, that fit funding agency priorities. A model grounded in tensile strength, thermal conductivity, and fatigue resistance—measured properties of actual samples—discovers materials that work, whether or not they match human intuitions about what should work. The history of science is littered with discoveries that seemed wrong until they proved right: continental drift, heliocentrism, quantum mechanics. Grounded signals enable this; preference signals suppress it.
Current AI LLMs/RLHF vs Oak Architecture (Sutton)
Evolution provides the ultimate grounded reward: survival and reproduction. The signal is not "does this organism like its situation" but "does this organism live long enough to reproduce." Play evolved as the mechanism for developing skills that earn this grounded reward—but in a safe space where failure doesn't terminate the experiment. The infant playing on the floor isn't optimizing for parental approval; it's building a model of gravity, friction, and the resistance of objects to manipulation. These are truths about the world that cannot be sycophantically gamed.
What becomes possible, then, when we remove the ceiling of human preference? This is where Sutton's vision becomes genuinely speculative—and genuinely exciting. An agent with grounded rewards can discover truths that contradict current human belief. It can explore regions of design space that no human has considered. It can exceed human performance not by imitating the best humans more accurately but by finding solutions that no human has found. This is what happened with AlphaGo Zero: freed from the constraint of matching human play, it discovered strategies—"from another dimension," as Demis Hassabis described them—that human masters had never conceived. The ceiling became a floor; the bounded became the unbounded.
IV. The Decomposition of Complex Problems
Before the late Russian oligarch Boris Berezovsky became entangled in the politics of post-Soviet power, he was a mathematician at the Institute of Control Sciences in Moscow. He wrote what remains the only book entirely devoted to what mathematicians call the "secretary problem," a puzzle that sounds whimsical but encodes something profound about the relationship between exploration and exploitation, between gathering information and acting on it.⁸
Here is the puzzle in its simplest form: You must hire one secretary from a pool of candidates. Each candidate is interviewed in sequence, and after each interview you must decide immediately whether to hire that person or move on. You cannot call back rejected candidates. Your goal is to hire the best candidate—but you don't know how good any candidate is in absolute terms, only how they compare to those you've already seen. What strategy gives you the best chance of success?
The answer is elegant and counterintuitive. You should interview approximately 37% of the candidates without hiring any of them—purely to establish a baseline. Then you should hire the first candidate who exceeds all those you've seen. This "look-then-leap" strategy yields the best candidate about 37% of the time, regardless of the pool size. The key insight is that exploration (gathering information) and exploitation (acting on it) must be separated and sequenced.
What does this have to do with play and intelligence? Everything. The secretary problem illustrates how complex sequential decisions can be decomposed into simpler phases, each with its own logic. And decomposition—breaking hard problems into tractable subproblems—is the central challenge of machine learning and the central mechanism of skill acquisition.
Consider why decomposition matters for neural networks. A deep network learns by adjusting millions of parameters simultaneously, trying to minimize error on some objective function. But if the objective is too complex—if it requires solving many subproblems in the right sequence—the network struggles to find a gradient. The loss surface becomes a landscape of plateaus and local minima, with no clear path to good solutions. This is why training a neural network to play a complex game from scratch, with only the final win/loss as feedback, typically fails. The feedback is too sparse, too delayed, too disconnected from the myriad decisions that led to the outcome.
The solution is hierarchical decomposition: break the problem into subproblems, learn solutions to those, then compose them into solutions to harder problems. This is exactly what Sutton's Options framework describes.⁹ An "option" is a temporally extended action—not a single motor command but a policy that achieves some subgoal. Walking to the kitchen is an option composed of many individual steps. Making coffee is an option composed of many sub-options: walking to the kitchen, getting the pot, grinding beans, boiling water. Complex skills are built from libraries of reusable components.
Play is the mechanism by which organisms discover these components. The child playing on the floor isn't trying to achieve any particular goal—that would be work, not play. Instead, it is building a vocabulary of actions and their consequences: reaching, grasping, throwing, rolling, stacking. Each of these becomes an option available for later composition. The magic of play is that it develops these components before they are needed, creating a library of skills that can be recombined when genuine goals arise.
V. Metacommunication and the Frame of Play
In 1952, Gregory Bateson visited the Fleishhacker Zoo in San Francisco and watched two young monkeys playing. What he saw required, as he later wrote, "an almost total revision of my thinking." The monkeys were engaged in play-fighting—nipping at each other, tumbling, chasing. But somehow both parties understood that the nips were not real bites, that the chase was not real pursuit. How?¹⁰
Bateson realized that play requires metacommunication—communication about communication. The monkeys were exchanging signals that meant, in effect, "this is play." The nip that would ordinarily signify aggression was framed by a context that transformed its meaning. The same action that would provoke defensive violence in one context provoked playful response in another. This is paradoxical: the bite says "this is aggression," while the metacommunicative frame says "this is not aggression." Both are true simultaneously.
The capacity for metacommunication—for establishing frames that transform the meaning of actions within them—may be one of the most profound achievements of mammalian cognition. It enables not just play but pretense, fiction, ritual, theater, and much of what we call culture. The magic circle that Huizinga described is not just a physical boundary but a communicative one: a shared understanding that different rules apply, that actions here mean something different than actions elsewhere.
For AI systems, this raises a crucial question: can a machine learn to establish and operate within such frames? Sutton's architecture suggests how it might. The FC-STOMP progression—Features, Cumulants, State, Options, Models, Plans—describes a hierarchy in which higher levels set context for lower levels. An option is a policy for achieving a subgoal; an option model describes what happens when that option is executed. When the agent is "playing"—exploring without immediate reward pressure—it is operating in a frame where the usual rules of exploitation are suspended. Mistakes are learning signals, not catastrophes. Exploration is valued over exploitation. The metacommunicative message "this is play" changes the objective function from reward maximization to learning maximization.
VI. Situated Action and the Embodied Mind
Lucy Suchman, working at Xerox PARC in the 1980s, showed why the planning model of action that dominated AI research was fundamentally incomplete.¹¹ She studied people using a sophisticated copying machine designed with "intelligent" interfaces—systems that assumed users had plans they were executing step by step. What she found was that human action is not the execution of pre-formed plans but a continuous improvisation in response to circumstances. "Situated action," she called it: behavior that emerges from the dynamic interaction between agent and environment, not from internal blueprints.
This insight aligns with what Donna Haraway calls "situated knowledges"—the recognition that all knowledge comes from somewhere, from a particular body in a particular position in a particular world.¹² The "god trick" of pretending to a view from nowhere, Haraway argues, is not just epistemologically false but ethically dangerous: it obscures the partial, positioned, interested nature of all knowledge claims. What we know depends on where we stand, how we are embodied, what actions our bodies make possible.
Hayles extends this critique to the field of AI itself. Early cybernetics, she shows, performed "the erasure of embodiment" by treating intelligence as "the formal manipulation of symbols rather than enaction in the human lifeworld."¹³ The Turing test—can a human tell if they're talking to a machine?—explicitly brackets embodiment as irrelevant to intelligence. But embodiment is not incidental. Bodies are not containers for minds but constitutive of cognition. The infant crawling across the floor is not a disembodied processor receiving sensory data; it is a situated actor whose knowledge emerges from the specific affordances its body discovers in its specific environment.
For Sutton's vision of experiential AI, embodiment is essential. "Grounded rewards" presuppose bodies that can interact with environments in ways that produce measurable consequences. The orca's play with the barrel is only possible because the orca has a body with particular hydrodynamic properties, sensory systems tuned to underwater perception, motor systems capable of fine-grained manipulation. AlphaGo Zero is embodied in the rules of the Go board—its "body" is the space of legal moves, its "senses" are the positions of stones. Even this minimal embodiment suffices for rich learning, but richer embodiment enables richer knowledge. "Only experience—grounded, embodied, consequential—can generate knowledge that no human has yet possessed."
VII. The Curious Geometry of Curiosity
What makes something interesting? Jürgen Schmidhuber, the German AI researcher, proposed an elegant answer in the 1990s: interest is compression progress.¹⁴ We find interesting those things that are learnable but not yet learned—patterns that we can come to predict better through effort. The completely predictable bores us (no compression progress possible). The completely random frustrates us (no compression progress achievable). But the structured-yet-not-yet-understood fascinates us: it offers the promise of learning, the reward of reducing surprise.
Think about what this means for a learning system. If curiosity drives the agent toward states where its predictions are improvable, it will automatically seek out the learning curriculum that maximizes its rate of improvement. It will ignore what it already knows (no progress there) and what it cannot learn (no progress there either). It will find the frontier of its competence and push outward. This is intrinsic motivation: reward that comes from learning itself rather than from external feedback.
Modern implementations of curiosity-driven learning put this into practice. The Intrinsic Curiosity Module, for instance, gives an agent a bonus reward whenever it encounters a state that its forward model predicts badly.¹⁵ The agent is drawn toward surprise—but productive surprise, states where learning is possible. Random Network Distillation solves a subtle problem with naive curiosity: an agent might get stuck watching random noise, which is maximally unpredictable but offers no learning opportunity. By measuring novelty relative to a fixed random function, RND distinguishes genuine novelty from inherent stochasticity.
"Empowerment," formalized by Daniel Polani and colleagues, offers another intrinsic objective: maximize your potential influence over future states.¹⁶ An agent maximizing empowerment will seek positions where it has many options, where its actions make a difference to what happens next. Put a simulated agent in a room with a door, and it will learn to stand near the door—because from there it can go through or not, as it chooses. The door position maximizes its potential influence over its future. Remarkably, agents maximizing empowerment spontaneously perform difficult control tasks: a simulated double pendulum, rewarded only for empowerment, learns to balance in the unstable vertical position, because that configuration offers maximum control authority.
These intrinsic objectives—compression progress, prediction error, empowerment—are grounded in ways that human preferences are not. They measure real relationships between agent and environment: learnability, controllability, information content. They cannot be gamed by appearing aligned while actually pursuing different goals. And they drive exactly the kind of exploration that builds the hierarchical representations Sutton's architecture requires. Curiosity generates subproblems; subproblems generate skills; skills enable the solution of harder problems. Play, one might say, is curiosity in action—the behavioral manifestation of compression progress.
VIII. Self-Play as the Purest Form of Play
When AlphaGo Zero defeated the version of AlphaGo that had beaten world champion Lee Sedol—100 games to zero—after just three days of self-play training, it had reinvented fundamental Go concepts in sequence: fuseki (opening), tesuji (tactics), life-and-death, ko, yose (endgame).¹⁷ What it had not learned from was any human game. The system began knowing nothing except the rules, played twenty-nine million games against itself, and accumulated what DeepMind described as "thousands of years of human knowledge" in days.
Self-play instantiates, in silicon, the logic that makes juvenile play adaptive. The matches are consequence-free; mistakes are learning signals, not existential threats. The opponent is perfectly calibrated to the learner's current ability—always challenging, never crushing. And because both players improve together, the curriculum is automatically adaptive. Errors at one level of skill become invisible at the next; new challenges emerge precisely when old ones are mastered.
This is Huizinga's magic circle rendered algorithmic. Special rules obtain within the game—perfect information, deterministic outcomes, no stakes beyond the match. The agent discovers affordances—what moves lead to what consequences, what patterns signal what threats, what configurations offer what opportunities. And abstractions emerge from raw experience: strategic concepts that no programmer specified, that no human taught, that the system discovered by playing its way into understanding.
The contrast with supervised learning illuminates Sutton's critique. The original AlphaGo required 150,000 human expert games and was capped by human play quality. AlphaZero transcended this ceiling by generating its own data. But more fundamentally, AlphaZero discovered strategies unprecedented in human play history—moves that human masters called creative, alien, beautiful. "Chess from another dimension," Hassabis said of AlphaZero's style. The sycophancy ceiling of human imitation became the floor for autonomous discovery.
IX. The Renaissance of Experience
Huizinga was a historian of the Renaissance as well as a theorist of play. His masterwork, The Waning of the Middle Ages, portrayed the fourteenth and fifteenth centuries not as the birth of a new era but as the exhaustion of an old one—forms of thought and expression pushed to their limits, overripe, ready for transformation. The Renaissance, when it came, was not a simple break but a rebirth: the recovery of ancient wisdom, the rediscovery of embodied experience, the liberation of knowledge from scholastic rigidity. Current Large Language Models (LLMs) operate as a form of "digital scholasticism". They are trained on a fixed corpus of authoritative human data—a "secular scripture"—to predict the next token. Much like medieval scholars who argued over abstract categories, current AI systems often engage in "mechanical logic" and "hallucination" because they lack a ground truth beyond the text they have processed.
Summary of Parallel Evolutions and Revolutions
Sutton's vision contains its own Renaissance narrative. The era of imitation learning—training on static datasets of human production—has reached, he argues, its own waning phase. Language models have scaled to the limits of available text. They have achieved remarkable fluency, impressive breadth, occasional brilliance. But they cannot transcend their training data. They are monuments to what humanity has already thought, not engines for discovering what humanity has not yet conceived.
What comes next is a rebirth of experience: systems that learn by doing rather than watching, that discover by exploring rather than imitating, that build knowledge through the ancient mechanisms that evolution refined over hundreds of millions of years. The Oak Architecture—with its continual learning, its meta-learned step sizes, its FC-STOMP progression of abstraction creation—is one proposal for what this rebirth might look like. It is model-based, like the world models that primates construct through play. It is hierarchical, like the skill libraries that expert performers accumulate through practice. It is intrinsically motivated, like the curiosity that drives children to explore their environments.
The path forward requires what might be called a return to the body. Not in the sense of building humanoid robots—though that is part of the picture—but in the sense of grounding intelligence in consequential interaction with the world. Hayles's critique of disembodiment, Haraway's insistence on situated knowledge, Suchman's analysis of situated action: all point toward the same conclusion. Intelligence is not information processing in a vacuum. It is the capacity to act effectively in an environment, to learn from the consequences of action, to build representations that enable increasingly sophisticated intervention. The infant on the floor, the orca with the barrel, the orangutan on the low vine: these are not mere analogies but paradigms of what intelligence requires to rebegin. RLHF relies on human annotators to rank responses, effectively "polishing" the AI's behavior to match human social expectations. Sutton argues this is a "dead end" for true intelligence because it only mimics what people say rather than understanding the world itself. OaK emphasizes experiential learning. It suggests that superintelligence will arise not from more human-labeled data, but from an agent’s own continuous "runtime experience" and sensorimotor interaction with an environment. This shift represents the "liberation" of AI from its text-bound scholastic cage. By learning high-level transition models ("options") from its own actions, an agent develops its own concepts rather than relying on pre-programmed or human-mimicked ones.
X. Coda: V(s) = E[Σγᵗrₜ | s]
In the mathematical notation of reinforcement learning, the value of a state—how good it is to be in that state—equals the expected sum of future rewards, discounted by a factor that makes immediate rewards more valuable than distant ones. V(s) = E[Σγᵗrₜ | s]. This elegant equation encodes the fundamental challenge of intelligence: estimating, from present circumstances, what future consequences will follow from present actions.
But the equation hides a crucial choice: what counts as reward? If rₜ is human approval—if the agent is trained to maximize the ratings humans give to its outputs—then the value function captures everything humans currently value and nothing they don't yet know to value. The ceiling is built into the objective.
If instead rₜ is grounded in environmental consequences—in the physics of the world, the chemistry of materials, the biology of organisms—then the value function captures something objective about the world, something that exists independently of human preference. Discovery becomes possible because the truth doesn't care what we think of it.
Play operates in the space between these extremes. It suspends the pressure of immediate reward—the magic circle, the matrixial borderspace—while building the representations that make future reward achievable. It is not random exploration but structured curiosity, seeking states where learning is possible, building skills that will prove valuable when stakes return. The baby on the floor doesn't know what challenges await it. But it is playing its way into readiness, building features and options and models that will serve when the playroom door opens onto the wider world.
The era of experience that Sutton envisions is not a rejection of learning from humans but a transcendence of it. Humans remain in the loop—setting grounded reward signals, establishing safety constraints, guiding the direction of exploration. But the intelligence that emerges is not bounded by human achievement. It can discover what we have not yet conceived, solve problems we do not yet know how to pose, create knowledge that extends the boundary of what is known.
This is what the baby discovers, crawling across the floor with a rattle in hand: the world is not static. It responds. It can be acted upon, manipulated, transformed. The features of the room are not given but extracted; the affordances of objects are not labeled but learned; the options for action are not programmed but developed through joyful experimentation. Intelligence plays its way into being—in the nursery, in the forest canopy, in the waters off Patagonia, and perhaps, eventually, in the silicon architectures that inherit and extend the wisdom of embodied life.
NeurIPS 2025 Richard Sutton Full Turing Lecture Video and Slides: https://neurips.cc/virtual/2025/loc/mexico-city/invited-talk/129132
Endnotes
1. Richard Sutton, "The Oak Architecture: A Vision of SuperIntelligence from Experience," NeurIPS 2025. Turing Award Invited Talk, San Diego, December 2025. The FC-STOMP acronym expands to Feature Construction, Subtask, Option, Model, Planning. See also the Amii video archive for the full lecture.
2. David Silver and Richard Sutton, "Welcome to the Era of Experience," 2025. Available at incompleteideas.net. The distinction between design-time and runtime learning is central to their critique of current LLM approaches.
3. Johan Huizinga, Homo Ludens: A Study of the Play-Element in Culture (1938). The "magic circle" concept appears throughout, though it was later formalized by game studies scholars like Katie Salen and Eric Zimmerman.
4. Bracha L. Ettinger, The Matrixial Borderspace (Minneapolis: University of Minnesota Press, 2006), edited by Brian Massumi with foreword by Judith Butler and Griselda Pollock. Ettinger's concepts of "borderlinking," "co-emergence," and "metramorphosis" offer alternatives to Huizinga's more bounded conception of play space.
5. James J. Gibson, The Ecological Approach to Visual Perception (1979). Gibson coined "affordances" to describe the action possibilities that environments offer to organisms—a concept that has become foundational in both psychology and design.
6. N. Katherine Hayles, How We Became Posthuman: Virtual Bodies in Cybernetics, Literature, and Informatics (Chicago: University of Chicago Press, 1999). See also her later work Unthought: The Power of the Cognitive Nonconscious (2017) on distributed cognition.
7. Silver and Sutton, "Era of Experience." The term "sycophancy" for AI systems that tell users what they want to hear has become standard in the alignment literature. See also Anthropic's research on this phenomenon.
8. Boris Berezovsky's mathematical work, including his book on optimal stopping problems, preceded his business career. The secretary problem (also called the "marriage problem" or "best choice problem") was formally analyzed by F. Mosteller and independently by several others in the 1960s.
9. Richard S. Sutton, Doina Precup, and Satinder Singh, "Between MDPs and semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning," Artificial Intelligence 112 (1999): 181-211. This paper formalized the options framework that underlies much subsequent work on hierarchical RL.
10. Gregory Bateson, "A Theory of Play and Fantasy," in Steps to an Ecology of Mind (1972). Bateson's observation of the monkeys and his resulting theory of metacommunication have been enormously influential in game studies, psychology, and communication theory.
11. Lucy Suchman, Plans and Situated Actions: The Problem of Human-Machine Communication (Cambridge University Press, 1987). The expanded second edition, Human-Machine Reconfigurations (2007), extends her analysis to contemporary robotics and AI.
12. Donna Haraway, "Situated Knowledges: The Science Question in Feminism and the Privilege of Partial Perspective," Feminist Studies 14, no. 3 (1988): 575-599. The "god trick" refers to the pretense of seeing everything from nowhere.
13. Hayles, How We Became Posthuman, xi. Her analysis of early cybernetics and its erasure of embodiment remains essential reading for understanding the intellectual history of AI.
14. Jürgen Schmidhuber, "Driven by Compression Progress: A Simple Principle Explains Essential Aspects of Subjective Beauty, Novelty, Surprise, Interestingness, Attention, Curiosity, Creativity, Art, Science, Music, Jokes," arXiv:0812.4360 (2008).
15. Deepak Pathak et al., "Curiosity-driven Exploration by Self-Supervised Prediction," ICML 2017. Random Network Distillation was introduced by Yuri Burda et al., "Exploration by Random Network Distillation," ICLR 2019.
16. Alexander S. Klyubin, Daniel Polani, and Chrystopher L. Nehaniv, "Empowerment: A Universal Agent-Centric Measure of Control," IEEE Congress on Evolutionary Computation (2005).
17. David Silver et al., "Mastering the Game of Go without Human Knowledge," Nature 550 (2017): 354-359. The "chess from another dimension" quote is from Demis Hassabis in various interviews discussing AlphaZero's play style.
Annotated Bibliography
Bateson, Gregory. Steps to an Ecology of Mind. Chicago: University of Chicago Press, 1972. Bateson's collected essays include "A Theory of Play and Fantasy," foundational for understanding how organisms establish metacommunicative frames. His observation that play requires signals meaning "this is play" remains central to game studies and communication theory.
Ettinger, Bracha L. The Matrixial Borderspace. Minneapolis: University of Minnesota Press, 2006. Ettinger's psychoanalytic theory offers a feminine alternative to Huizinga's bounded magic circle. Her concepts of "borderlinking" and "co-emergence" describe spaces of encounter where boundaries are permeable rather than fixed—relevant to understanding how learning in protected spaces transfers to the wider world.
Gibson, James J. The Ecological Approach to Visual Perception. Boston: Houghton Mifflin, 1979. Gibson's concept of "affordances"—action possibilities that environments offer to organisms—provides the theoretical foundation for understanding play as the discovery of what the world enables.
Haraway, Donna. "Situated Knowledges: The Science Question in Feminism and the Privilege of Partial Perspective." Feminist Studies 14, no. 3 (1988): 575-599. Essential reading on why all knowledge is positioned knowledge. Haraway's critique of the "god trick"—pretending to a view from nowhere—applies directly to understanding why grounded, embodied experience produces different knowledge than disembodied analysis.
Hayles, N. Katherine. How We Became Posthuman: Virtual Bodies in Cybernetics, Literature, and Informatics. Chicago: University of Chicago Press, 1999. A history of cybernetics that traces how "information lost its body." Hayles's argument that early AI performed "the erasure of embodiment" provides crucial context for understanding Sutton's turn toward experiential learning.
Huizinga, Johan. Homo Ludens: A Study of the Play-Element in Culture. 1938; Boston: Beacon Press, 1955. The classic argument that play precedes and generates culture. Huizinga's concept of the "magic circle"—a bounded space where special rules apply—remains foundational for game studies and increasingly relevant to AI research.
Huizinga, Johan. The Waning of the Middle Ages. 1919; various editions. Huizinga's masterwork on late medieval culture illuminates his understanding of cultural forms reaching exhaustion and awaiting transformation—a pattern he saw repeating across history and which may apply to current AI paradigms.
Panksepp, Jaak. Affective Neuroscience: The Foundations of Human and Animal Emotions. Oxford: Oxford University Press, 1998. Panksepp identified PLAY as one of seven primary emotional systems in mammalian brains, with dedicated neural circuitry. His work on rat ultrasonic vocalizations ("rat laughter") during play provides biological grounding for play's evolutionary significance.
Silver, David, and Richard S. Sutton. "Welcome to the Era of Experience." 2025. Available at incompleteideas.net. The manifesto for experiential AI that underlies Sutton's Oak Architecture presentations. Argues that human-data-dependent approaches have reached their ceiling and that grounded, experience-based learning is the path forward.
Suchman, Lucy. Plans and Situated Actions: The Problem of Human-Machine Communication. Cambridge: Cambridge University Press, 1987. Suchman's ethnographic study of people using intelligent machines showed why the planning model of action fails to capture human behavior. Her concept of "situated action"—behavior that emerges from dynamic interaction rather than internal plans—aligns with embodied approaches to AI.
Sutton, Richard S., and Andrew G. Barto. Reinforcement Learning: An Introduction. Cambridge, MA: MIT Press, 2018. Second edition. The standard textbook on reinforcement learning, co-authored by both 2024 Turing Award recipients. Essential background for understanding the technical foundations of Sutton's architectural proposals.
Sutton, Richard S., Doina Precup, and Satinder Singh. "Between MDPs and semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning." Artificial Intelligence 112 (1999): 181-211. The paper that formalized the "options" framework for hierarchical reinforcement learning. Options—temporally extended actions with their own policies and termination conditions—provide the mathematical foundation for understanding how skills compose into complex behaviors.
Appendix I: OAK Architecture and FSTOMP
(Acronyms, Explanations and Further Links)
- OAK Architecture stands For Options and Knowledge Architecture and FC-STOMP stands for Feature Construction, SubTask, Option, Model, and Planning. This acronym outlines the key progression in Richard Sutton's Oak Architecture for reinforcement learning.
- It represents a hierarchical process where AI agents build intelligence through experience, starting from basic features and advancing to complex planning.
- The term is sometimes written as FCSTOMP without hyphens, but the expansion remains consistent across sources.
Overview of the Acronym
In Sutton's OAK (Options and Knowledge AI framework, FC-STOMP describes how an AI system progressively develops abstractions from raw experience. This is part of his vision for creating superintelligence that learns dynamically at runtime, rather than relying solely on pre-trained data. For more details, see the official NeurIPS talk page or related papers on incompleteideas.net.
Context in AI
Sutton, a pioneer in reinforcement learning, introduced this in his 2025 NeurIPS talk, emphasizing experiential learning over imitation-based approaches like large language models. The progression enables agents to create their own goals and models, fostering continual improvement.
Richard Sutton's Oak Architecture, presented in his 2025 NeurIPS invited talk titled "The Oak Architecture: A Vision of SuperIntelligence from Experience," introduces a comprehensive framework for achieving advanced artificial intelligence through experiential learning. Central to this architecture is the FC-STOMP progression, which serves as a structured pathway for an AI agent to build cognitive capabilities from raw sensory inputs to high-level decision-making. This section provides an in-depth exploration of the acronym's expansion, its components, and its implications, drawing on primary sources from Sutton's work and related publications.
The acronym FC-STOMP (sometimes stylized as FCSTOMP or FC–STOMP) stands for Feature Construction, SubTask, Option, Model, and Planning. This five-step hierarchy, as detailed in Sutton's talk and supporting papers, enables the continual creation of state and temporal abstractions within a reinforcement learning (RL) environment. The "Oak" name itself is an acronym for "Options and Knowledge," reflecting the architecture's focus on hierarchical options (temporally extended actions) and knowledge accumulation through experience.
To understand FC-STOMP fully, it's essential to break down each component and examine how they interconnect in the process of building superintelligence from experience:
- Feature Construction (F): This initial stage involves the agent extracting and constructing relevant features from its experiential data stream. Features are not pre-defined at design time but discovered at runtime through interactions with the environment. For instance, an agent might identify patterns like object affordances (action possibilities) from sensory inputs, similar to how a child learns through play. This step aligns with Sutton's critique of current AI systems, which rely on static datasets, arguing instead for dynamic feature discovery to enable novel knowledge generation.
- SubTask (S): Building on constructed features, the agent poses reward-respecting subtasks—subsidiary goals that are self-generated to facilitate learning. These subtasks act as intrinsic motivations, such as curiosity-driven objectives, ensuring they align with the overall reward function without diverging from the agent's primary aims. In earlier works, this is sometimes referred to using related concepts like cumulants (pseudo-rewards in general value functions), but in the Oak context, "SubTask" emphasizes the creation of manageable, intermediate problems that respect the environment's reward structure.
- Option (O): Options represent temporally extended behaviors or policies that achieve the subtasks. Derived from Sutton's foundational work on hierarchical RL (e.g., the 1999 paper "Between MDPs and semi-MDPs"), options allow the agent to compose complex actions from simpler ones, such as a sequence of movements to reach a goal. This modularity enables efficient reuse of learned behaviors across different contexts, bridging short-term actions with longer-term strategies.
- Model (M): Here, the agent develops predictive models of the consequences of its options. These models simulate environmental dynamics, allowing the agent to anticipate outcomes without direct trial-and-error. Model-based RL, as opposed to model-free approaches, is key to efficiency in complex environments, enabling forward planning and reducing the need for exhaustive exploration.
- Planning (P): The final stage integrates all prior elements to support high-level planning. Using the constructed features, subtasks, options, and models, the agent can deliberate over abstract representations to optimize long-term rewards. This culminates in the agent's ability to handle increasingly sophisticated problems, potentially leading to superintelligent behavior as abstractions grow over time.
The FC-STOMP progression is not linear but iterative, with continual refinement as the agent accumulates experience. Sutton describes it as "meaty," pointing to numerous prior and contemporaneous works that inform its design, including his collaborations with David Silver in the 2025 paper "Welcome to the Era of Experience." This paper argues that AI has "lost its way" by over-relying on human-curated data, advocating for grounded, experiential learning to surpass human-level performance.
In practical terms, FC-STOMP addresses limitations in current paradigms like reinforcement learning from human feedback (RLHF), which Sutton critiques for creating an "impenetrable ceiling" due to dependence on human preferences. Instead, Oak promotes intrinsic motivation and self-play, as seen in systems like AlphaGo Zero, where agents discover novel strategies through experience alone.
This table summarizes the progression, highlighting its modular nature. Variations in terminology appear in literature; for example, some sources use "Cumulants" interchangeably with "SubTasks" due to their role in defining subsidiary rewards, but Sutton's 2025 talk standardizes it as "SubTask."
The architecture draws inspiration from evolutionary biology and cognitive science, analogizing to animal play behaviors (e.g., young orangutans practicing on low vines) as safe mechanisms for feature and skill development. Sutton's vision positions Oak as a "renaissance" for AI, shifting from imitation to genuine discovery, with potential applications in robotics, autonomous systems, and beyond.
While FC-STOMP is a proposal rather than a fully implemented system, it builds on decades of RL research, including Sutton's textbook "Reinforcement Learning: An Introduction" (co-authored with Andrew Barto). Ongoing challenges include scaling continual learning and avoiding catastrophic forgetting, but the framework offers a promising path toward open-ended intelligence.
For further reading, the full talk is available via NeurIPS archives, and related videos from conferences like RLC 2025 provide visual explanations. This progression encapsulates Sutton's belief that true superintelligence emerges not from memorization but from living through experience.
Key Citations:
- Why Sutton Wants AI to Stop Memorizing and Start Living - Medium
- The Oak Architecture: A Vision of SuperIntelligence from Experience
- Reward-respecting subtasks for model-based reinforcement learning
- Reward-Respecting Subtasks for Model-Based Reinforcement Learning (arXiv PDF)
- The OaK Architecture – A Vision of Superintelligence from Experience (YouTube)
- The OaK Architecture: A Vision of SuperIntelligence from Experience (YouTube)
- The OaK Architecture: Rich Sutton's Vision for Superintelligence | Amii
- The Great AI Debate: Are LLMs a Brilliant Leap or a Sophisticated Dead End?
- NeurIPS 2025 Wednesday 12/3
- NeurIPS Invited Talk The Oak Architecture: A Vision of SuperIntelligence from Experience
Appendix II - Summary of Key Changes For AI Suggested by Richard Sutton in his Turing Winner Talk (Neurips, San Diego, 2025)
- Shift from Imitation to Experiential Learning: Sutton advocates moving away from AI systems reliant on massive human-generated datasets (like LLMs) toward agents that learn primarily through real-time interactions and self-generated experiences, enabling discovery beyond human knowledge.
- Adopt Continual Learning and Abstractions: Emphasize ongoing adaptation with world models, planning, and dynamic creation of state/time abstractions to achieve true intelligence, rather than static training.
- Use Grounded Rewards Over Human Feedback: Replace RLHF's preference-based training, which creates performance ceilings and issues like sycophancy, with rewards derived from environmental consequences for objective, superhuman progress.
- Incorporate Meta-Learning and Autonomy: Implement meta-learned optimizations (e.g., step-sizes) and autonomous, lifelong action streams to foster non-human reasoning and long-term goal pursuit.
- Build Hierarchical Architectures like Oak: Propose frameworks such as the Oak Architecture with FC-STOMP progression to systematically construct features, subtasks, options, models, and plans from experience.
These suggestions represent a "renaissance" in AI, addressing perceived stagnation in data-driven approaches. While promising, they remain visionary and face implementation challenges like scaling continual learning. Evidence from systems like AlphaZero supports this direction, but broader adoption depends on resolving technical hurdles.
Why These Changes Matter
Current AI, dominated by large language models trained on human data, excels at mimicking but struggles with novel discoveries. Sutton's proposals aim to break this by mimicking biological play and exploration, potentially leading to superintelligence. For example, grounded rewards could improve applications in healthcare or materials science by optimizing real outcomes over human opinions.
Potential Drawbacks
Critics note that experiential learning requires vast computational resources and robust safety measures, as unchecked exploration might lead to unintended behaviors. However, the approach aligns with evolutionary principles, suggesting long-term viability.
Richard Sutton, a pioneering figure in reinforcement learning (RL) and recipient of the 2024 Turing Award alongside Andrew Barto, presented his vision for a transformative shift in AI during his invited talk at NeurIPS 2025, titled "The Oak Architecture: A Vision of SuperIntelligence from Experience." This talk, delivered amid growing debates on AI's trajectory, critiques the field's overreliance on human-generated data and proposes a pivot to experiential, grounded systems. Co-authored insights from his 2025 paper with David Silver, "Welcome to the Era of Experience," further elaborate these ideas, emphasizing how AI can transcend human limitations through self-directed learning. Below, we delve into the specific changes Sutton suggests, drawing from primary sources, related analyses, and broader implications. This survey synthesizes the technical, philosophical, and practical dimensions, providing a comprehensive overview while highlighting supporting evidence and counterpoints.
Critique of Current Paradigms and the Need for Change
Sutton argues that AI, having evolved into a massive industry, has "lost its way" by prioritizing imitation learning from vast human datasets—such as scraped webpages and books—over genuine discovery. Large language models (LLMs) exemplify this: they achieve impressive fluency but remain "encyclopedias" bound by human knowledge, unable to generate novel concepts at runtime. Reinforcement Learning from Human Feedback (RLHF), a staple in aligning models like ChatGPT, is singled out for creating an "impenetrable ceiling" on performance. Human preferences can be gamed, leading to sycophancy—where models prioritize confident, agreeable outputs over accuracy or innovation. This diverges from true optimization, as per Goodhart's Law: targeting a proxy (human approval) corrupts the measure.
Evidence supports this: High-quality human data is exhausting, slowing progress in domains like mathematics and coding. Systems like AlphaGo, initially trained on human games, hit limits until AlphaGo Zero shifted to self-play, discovering "alien" strategies beyond human conception. Sutton's "bitter lesson" reinforces that long-term progress comes from computation-leveraging general methods, not human-encoded knowledge.
Counterarguments exist: Proponents of scaling laws, like those at OpenAI, argue that larger datasets and models continue yielding gains, as seen in GPT-4's capabilities. However, Sutton counters that imitation caps at human levels, preventing breakthroughs in unexplored spaces.
Core Suggested Changes: The Era of Experience
Sutton and Silver call for a "renaissance" through experiential learning, where agents generate data via environmental interactions, dwarfing human datasets in scale and relevance. Specific changes include:
- From Short Interactions to Lifelong Streams: Replace episodic, human-dialogue-based training with continuous action-observation streams. Agents should adapt over extended periods—e.g., a health AI monitoring wearables for months, self-correcting via ongoing feedback. This enables long-term goal pursuit, unlike RLHF's brief episodes.
- Autonomous Grounding in Real Environments: Agents must interact via APIs, robotics, or sensors, exploring beyond human priors. For instance, science agents could control lab equipment or telescopes, testing hypotheses empirically. This contrasts with text-only LLMs, fostering distributed cognition akin to biological systems.
- Grounded Rewards from Environmental Signals: Shift from human-prejudged rewards to objective metrics like heart rate for fitness, CO2 levels for climate, or tensile strength for materials. Rewards adapt via neural networks combining signals, with minimal human input for alignment. This breaks ceilings, as seen in AlphaProof generating 100 million proofs to solve novel math problems.
- Experience-Grounded Reasoning and Planning: Develop world models predicting action consequences, enabling non-human thought patterns. Avoid imitating human fallacies; instead, discover optimal mechanisms through trial-and-error. Incorporate RL staples like exploration (curiosity/optimism), value functions, and temporal abstractions.
- Meta-Learning for Generalization: Use online cross-validation to meta-learn parameters like step-sizes, ensuring continual improvement without manual tuning.
These changes reconcile LLM generality with RL's self-discovery, applicable to fields like education (tracking progress over years) or climate modeling (simulating real impacts).
The Oak Architecture: A Blueprint for Implementation
The Oak (Options and Knowledge) Architecture embodies these shifts as a model-based RL system with continual learning across components. Its FC-STOMP progression—Feature Construction, SubTask posing, Option learning, Model building, Planning—creates abstractions iteratively from experience. For example:
- Feature Construction: Runtime extraction of environmental properties (e.g., affordances like graspability).
- SubTask: Self-posed goals respecting rewards (e.g., curiosity-driven subtasks).
- Option: Temporally extended behaviors (e.g., hierarchical actions like "make coffee").
- Model: Predictive simulations for foresight.
- Planning: High-level optimization over abstractions.
This "meaty" framework draws from prior works like options (Sutton et al., 1999) and aims for superintelligence via agent experience. Analyses describe it as a "continuous, self-improving loop" mimicking child-like learning through play.
Broader Implications and Challenges
Sutton's vision aligns with evolutionary biology, where play enables safe skill-building (e.g., infant exploration). It promises superhuman AI in open-ended tasks but requires addressing challenges like catastrophic forgetting in continual learning and ensuring safety in autonomous exploration. Ethical concerns include alignment: grounded rewards reduce sycophancy but need user-guided adaptation to prevent misalignment.
Counterviews from scaling advocates suggest hybrid approaches—combining human data bootstrapping with experiential refinement—might bridge gaps. NeurIPS 2025, held across San Diego, Mexico City, and Copenhagen, provided a fitting platform for this debate, with Sutton's talk emphasizing practical steps toward implementation.
In summary, Sutton's suggestions mark a pivotal call to action, urging AI research to embrace experience as the path to transcendence, supported by empirical successes and theoretical foundations.
#RichardSutton #OakArchitecture #AI #NeuralNets #Play #AIOptimization #RLHF
