Raymond UzwyshynIdeas · Research · Artificial Intelligence
Science, Research & Discovery

AI and the Semiotic Languages of the Sciences

Jensen Huang (Nvidia): 2025 Cambridge UK Stephen Hawking Prize Talk: https://www.youtube.com/watch?v=YNshj2oOr3E (November 4, 2025

Cover graphic for AI and the Semiotic Languages of the Sciences

Ray Uzwyshyn, Ph.D. MBA MLIS

Jensen Huang (Nvidia): 2025 Cambridge UK Stephen Hawking Prize Talk: https://www.youtube.com/watch?v=YNshj2oOr3E (November 4, 2025)

1. Introduction: The Shift from Observation to Dialogue

The history of scientific inquiry has been dominated by a singular mode of engagement: observation. From the earliest taxonomists cataloging flora to the high-energy physicists smashing atoms in colliders, the scientific method has fundamentally relied on the human subject observing the material object. We measure, we record, and we attempt to deduce the silent laws governing matter from the outside looking in. However, we are currently standing at the precipice of a profound epistemological rupture, a shift so fundamental that it redefines the very nature of discovery. We are moving from an era of observation to an era of dialogue. We are no longer merely analyzing matter; we are entering into a conversation with it.

This paradigm shift was crystallized by Jensen Huang, CEO of NVIDIA, during his acceptance of the Stephen Hawking Prize at the Cambridge Union. In a moment of striking clarity, Huang posited that biology is transitioning from a field of sporadic scientific discovery to a field of engineering, largely because humanity has finally learned to "represent the language of biology".1 His assertion that artificial intelligence (AI) will allow scientists to "talk to proteins, asking them what they are and how they behave, just like we already do with images" is not merely a poetic metaphor. It is a precise, technical description of the capabilities of Large Language Models (LLMs) and Transformer architectures when applied to the fundamental substrates of life. By treating amino acids, atomic coordinates, and molecular functions as tokens in a high-dimensional semiotic system, AI has emerged as a universal translator, rendering the opaque languages of nature legible to human query.

This report provides an exhaustive analysis of this transition, arguing that the application of AI to natural sciences represents a "semiotic turn" in engineering. It explores the theoretical underpinnings of "biosemiotics" in the age of deep learning, examining how the structuralist theories of Claude Lévi-Strauss, Roland Barthes, and Umberto Eco are finding unexpected empirical validation in the architecture of neural networks. It details the technological mechanisms—specifically the Transformer architecture—that enable this translation, transforming the "slush" of biological data into structured, grammatical syntax. Furthermore, it extends this analysis beyond biology, demonstrating how similar semiotic principles are being applied to materials science, climate physics, and astrophysics, effectively positioning AI as a universal interpreter of material systems.

1.1 The Engineering of Biology

The distinction Huang draws between "science" and "engineering" is critical to understanding the current historical moment. Science is characterized by the exploration of the unknown, often relying on trial and error, serendipity, and sporadic breakthroughs. It is a process of uncovering what already exists. Engineering, conversely, relies on predictable principles, repeatable processes, and the compounding of benefits over time.1 For biology to transition from a descriptive science to a generative engineering discipline, it requires a representational system—a language—that is consistent, computable, and generative.

For decades, the pharmaceutical industry has operated on a model of "drug discovery," a term Huang derides as akin to "looking for mushrooms"—a process where researchers venture into the vast forest of chemical possibilities hoping to stumble upon a useful compound.5 This method is fraught with failure; it is sporadic, chaotic, and fundamentally unscalable. The digitization of biology changes this dynamic from discovery to design. By training AI models on the vast corpus of known protein structures and sequences—the "literature" of biology—we can now "write" new biological sentences. These are proteins that have never existed in nature but follow its grammatical rules to perform specific, engineered functions.4

1.2 The Semiotic Proposition

The central thesis of this report is that AI functions as a semiotic engine. Semiotics, the study of signs and symbols and their interpretation, provides the necessary intellectual framework for understanding how an algorithm designed to predict the next word in a sentence can also predict the 3D structure of a protein or the stability of a crystal lattice.

In the human conceptual system, a word is a signifier that points to a signified concept. In the biological system, an amino acid sequence (the signifier) determines a 3D structure (the signified), which in turn dictates biological function (the meaning).7 AI models like AlphaFold, ESM3, and GNoME have demonstrated that these material systems possess their own vocabularies, syntaxes, and grammars. By learning these hidden structures, AI allows us to translate human intent (e.g., "design a binder for this spike protein") into biological syntax (a specific sequence of amino acids). This capability represents a profound expansion of the "semiosphere," extending the realm of interpretable signs beyond human culture and into the very fabric of matter itself.9

1.3 The Platform Shift

This transition is not occurring in a vacuum but is underpinned by a massive platform shift in computing. Just as the transistor sparked the age of semiconductors and the internet connected the world's knowledge, the convergence of accelerated computing (GPUs) and generative AI is creating a new "industrial revolution".11 This shift is characterized by the movement from "retrieval" based computing (finding existing files) to "generative" computing (creating new data). In the context of biology, this means we are no longer limited to retrieving the biological solutions evolution has produced over billions of years; we can now generate the solutions evolution missed. The infrastructure supporting this—what Huang calls "AI factories"—is being built to process the exabytes of biological data required to train these models, turning the "secrets of biology into answers".11


2. The Theoretical Framework: Biosemiotics and Structuralism

To fully grasp the implications of "talking to proteins," one must navigate the deep intersection of linguistic theory, philosophy, and molecular biology. The application of language models to biology is not an arbitrary borrowing of tools from computer science; it rests on the fundamental ontological reality that life is, at its core, information processing.

2.1 Structuralism 2.0?

The resurgence of structuralism—the mid-20th-century intellectual movement that viewed culture, mythology, and language as systems governed by underlying structures—appears imminent in the age of AlphaFold. Structuralists like Claude Lévi-Strauss and Roland Barthes argued that diverse phenomena could be understood as languages with their own rules. Lévi-Strauss showed that cultural codes operate as structured languages; Barthes revealed how objects and images signify; Eco systematized how different domains function as sign systems with their own vocabularies, grammars, and logic.13

Today, computational biologists use strikingly similar rhetoric, often inadvertently. When Ali Madani of Profluent states that AI models learn from sequences "whether those are sequences of characters or words or computer code or amino acids," he is articulating a form of meta-structuralism.15 This view posits a universal isomorphism between different domains of reality, accessible through the formalism of sequence analysis.

Scholars like Fabian Offert have termed this development "Structuralism 2.0," noting the intuitive appeal of viewing proteins as having a "universal grammar" similar to human language.8 However, this analogy is complex and fraught with epistemic risk. While human language is discrete, symbolic, and socially constructed, biological "language" is often continuous, physical, and thermodynamically determined. The "meaning" of a protein is not a social convention but a physical reality—the folding of a chain into a low-energy state.16

2.2 Reverse Structuralism: The Epistemic Detour

Critically, the "structuralism" observed in AI protein folding is arguably a "reverse structuralism".8 We did not start with a comprehensive linguistic theory of biology and build a machine to test it. Rather, we built a machine (the Transformer) designed for natural language, applied it to biology due to data similarities, and discovered, somewhat surprisingly, that it worked. The structure of biological information was revealed post hoc through the empirical success of the tool.

This phenomenon suggests that the Transformer architecture captures a form of universality that transcends specific domains, modeling a "non-linguistic" continuity that underpins both text and chemistry.8 The tool itself—the Transformer—becomes a "self-defining epistemic object".8 It defines a specific class of knowledge (sequence-dependent, attention-based relationships) that turns out to be applicable to both Shakespearean sonnets and viral capsids. This "reverse structuralism" implies that we are discovering the structure of reality by seeing what our language machines can successfully simulate.

2.3 Biosemiotics: The Sign Systems of Life

Biosemiotics is the study of the production and interpretation of signs and codes in the biological realm. It posits that meaning-making is not exclusive to humans but is a fundamental property of all living systems.17 From a single cell navigating a chemical gradient to a protein binding to a receptor, biological entities are constantly interpreting environmental cues—signs.

Jensen Huang’s vision of "digital biology" aligns with the biosemiotic view that proteins operate within a system of signification.

  • Syntax: The linear sequence of amino acids ($A, C, D, E, F...$) constitutes the primary text.
  • Grammar: The physiochemical rules (hydrophobicity, electrostatic charge, van der Waals forces) act as the grammar, governing how these sequences can and cannot fold.
  • Semantics: The biological function (catalysis, binding, structural support) that emerges from the folded structure is the "meaning" of the sequence.7

Before the advent of AI, humans could read the syntax (sequencing DNA was solved decades ago) but could not reliably interpret the grammar to understand the semantics. We possessed the "book of life" but lacked the dictionary to translate it. AI serves as this interpreter. It has "read" the entire Protein Data Bank (PDB) and internalized the high-dimensional relationships between sequence and structure, allowing it to predict the "meaning" (structure/function) of a sequence with unprecedented accuracy.3

2.4 Not Minds, but Signs

It is crucial to clarify the ontological status of the AI in this arrangement. As Davide Picca argues in his paper "Not Minds, but Signs," these systems should not be framed as "cognitive systems" or "digital minds".9 They are "semiotic machines." They do not possess consciousness, intent, or "understanding" in the human sense. Instead, they are engines of symbolic recombination. They operate by determining the statistical probability of a token (word or amino acid) appearing in a specific context.

When an AI "talks" to a protein, it is not engaging in a cognitive conversation. It is navigating a latent space—a high-dimensional mathematical representation of all possible biological configurations. The "dialogue" is a mathematical traversal of this space. When a researcher prompts a model to "evolve this protein to withstand higher temperatures," the model calculates a vector path from the current sequence to a region of the latent space associated with thermal stability, and then decodes that vector back into a sequence of amino acids.10 This is semiotic manipulation at a scale and speed impossible for the human mind, effectively automating the "interpretation" of biological signs.


3. The Mechanism of Translation: Transformers in Biology

The technological engine driving this semiotic translation is the Transformer architecture, originally developed for Natural Language Processing (NLP) by Google researchers in 2017. Understanding why an architecture designed for English and French works for Proteins and RNA is key to understanding the "universal translator" concept.

3.1 Tokenization: The Alphabet of Matter

The first step in any language model is tokenization—breaking down complex data into discrete units. In NLP, tokens are words or sub-words. In digital biology, the tokens are the monomeric building blocks of life.

Article content

By reducing proteins to strings of tokens (e.g., M-A-G-H...), scientists can feed biological data into the same neural networks used for ChatGPT. However, unlike words, amino acids interact physically. A "word" at the beginning of a protein sequence might chemically bond with a "word" at the end of the sequence if the protein folds back on itself. This "long-range dependency" is exactly what Transformers were designed to solve in linguistics (e.g., resolving how a pronoun in the last sentence refers to a noun in the first).18

3.2 Attention Mechanisms as Physical Forces

The core innovation of the Transformer is the "self-attention" mechanism. It allows the model to weigh the importance of every token relative to every other token in the sequence. In a linguistic sentence, attention helps the model understand that "bank" refers to a river and not money based on the surrounding context words like "water" or "flow."

In a protein, the attention maps generated by the AI often correspond to actual physical contacts and forces. Research has shown that the "attention heads" in protein language models effectively learn the 3D contact map of the protein without being explicitly told the laws of physics.19 The AI learns that Residue 5 "pays attention" to Residue 150 because they form a hydrogen bond in the folded structure. Thus, the "grammar" the model learns is actually the physics of protein folding. This is a profound realization: the statistical patterns of the sequence encode the physical laws of the structure. The AI reads the sequence and "hallucinates" the physics.16

This challenges the traditional view that simulation must be based on "first principles" (physics-based modeling). Instead, AI demonstrates that physics can be learned phenomenologically from data. The attention mechanism effectively "rediscovers" the forces of nature by observing their effects on sequence statistics.

3.3 The Latent Space: A Map of the Possible

When a model like ESM3 (Evolutionary Scale Modeling) processes a protein, it compresses the information into an embedding—a dense vector in a high-dimensional space (often thousands of dimensions). In this latent space, proteins with similar functions are clustered together, even if their sequences look very different.21

This latent space is the modern "Rosetta Stone." It allows for translation between modalities.

  • Sequence-to-Structure: AlphaFold predicts the 3D coordinates from the 1D sequence by mapping the sequence embedding to structural coordinates.22
  • Structure-to-Sequence: Inverse folding models (like ProteinMPNN or ESM-IF) look at a desired 3D shape and predict the amino acid sequence needed to create it.23
  • Function-to-Sequence: Generative models can take a functional description (e.g., "binds to EGFR") and generate a sequence.

This multidimensional map enables "zero-shot" predictions—the ability to predict properties for proteins the model has never seen before, simply by analyzing their position in the latent space relative to known proteins.21 It turns biology into a geometry problem: finding a cure becomes a matter of finding the right vector.


4. The Language of Proteins: A Deep Dive

To understand how AI talks to proteins, we must look closer at the specific "linguistic" components of biological systems and how they are modeled.

4.1 Syntax: The Sequence of Life

The syntax of proteins is linear but information-dense. Nature uses a 20-character alphabet (amino acids), but unlike human alphabets, these characters have distinct physicochemical personalities. Tryptophan is bulky and hydrophobic; Lysine is positively charged and hydrophilic. The "spelling" of a protein determines its fate.

Language models trained on these sequences (like ESM-2 or ProtBERT) use "Masked Language Modeling" (MLM) to learn this syntax.24 The model is shown a sequence with some amino acids hidden (masked) and must guess what they are. By doing this billions of times, it learns the evolutionary constraints of protein sequences. It learns, for example, that a hydrophobic residue is rarely followed by a charged residue in a buried helix. This training process is analogous to a human learning English by filling in missing words in millions of sentences.

4.2 Grammar: Folding as Logic

If the sequence is the syntax, the folding is the grammar. The "grammar" of proteins is defined by Levinthal's paradox: a protein has an astronomical number of possible configurations, yet it folds into its native state in milliseconds. It follows a "grammatical" path to stability.

DeepMind's AlphaFold 2 and 3 cracked this grammatical code not by simulating the folding atom-by-atom (which is too slow), but by recognizing the patterns of the final folded state.22 AlphaFold 3 extends this to interactions, predicting how proteins interact with DNA, RNA, and ligands (drugs).27 It predicts the "sentence structure" of molecular complexes.

Interestingly, researchers have found that these models also work on disordered proteins. "Dr. BERT" is a model specifically designed to predict Disordered Regions (DR) in proteins—segments that do not have a fixed structure but are essential for signaling.28 This suggests that the "language" of biology includes "slang" or "free verse"—regions where the strict grammatical rules of folding are relaxed to allow for flexibility.

4.3 Semantics: Function and Meaning

In linguistics, semantics is the study of meaning. In biology, meaning is function. A protein "means" kinase activity, or oxygen transport, or viral capsid formation.

The breakthrough of recent models like ESM3 is the integration of semantics directly into the generation process. ESM3 is a "multimodal" model that reasons over sequence, structure, and function simultaneously.29 It can be prompted with a function (semantics) to generate a sequence (syntax).

Article content

5. The Dialogue: Talking to Proteins

"Talking to proteins" implies a bidirectional exchange. It is not enough to simply predict (read); we must also be able to design (write). This is where Generative AI enters the picture, transforming biology from an analytical science into a creative engineering discipline.

5.1 Human-to-Protein Conversation: The Prompt Interface

The interface for this new biology is increasingly resembling a chatbot. NVIDIA's BioNeMo framework and services like ESM3 allow researchers to "prompt" biology using natural language or specialized tags, much like generating an image with Midjourney.30

The Grammar of Biological Prompts:

Recent advancements allow for multi-modal prompting. A researcher can prompt a model with:

  1. Natural Language: "Generate a protein that binds to the SARS-CoV-2 spike protein."
  2. Structure: "Complete this protein backbone but change the active site loop" (Structure-based in-painting).
  3. Chemical Properties: "Design a sequence with high thermal stability and low viscosity."

ESM3, for example, acts as a generative masked language model. It treats sequence, structure, and function as "tracks" of information. One can mask the sequence track and provide the structure track, asking the model to "fill in the blanks" (inverse folding). Or one can provide a partial sequence and ask the model to auto-complete it (evolutionary extrapolation).29 This "fill-in-the-blank" capability is the direct functional equivalent of a conversation: the human provides the context, and the AI provides the response.

5.2 Case Study: The Chain of Thought in Fluorescent Proteins

A striking example of this dialogue is the work by EvolutionaryScale using ESM3 to generate a novel fluorescent protein. The researchers did not simply ask for a final product; they engaged the model in a "chain of thought" process.29

They prompted the model to reason over the sequence, structure, and function iteratively. They asked the model to generate a protein with specific fluorescence properties (function) and a specific structural fold (grammar). The result was a bright fluorescent protein that shared only 58% sequence identity with naturally occurring fluorescent proteins.24

To put this in perspective: in the world of biology, a 58% difference is massive. It is roughly equivalent to the divergence between species separated by 500 million years of evolution. If this were a language, it would be like the AI inventing a new dialect of English that is mutually intelligible with modern English but uses entirely different vocabulary roots. The model did not memorize a fluorescent protein; it understood the concept of fluorescence and the grammar of the beta-barrel fold well enough to write an entirely original "poem" that glowed green.

5.3 Protein-to-Protein Conversation: Eavesdropping on the Cell

The dialogue is not limited to Human-AI interactions. Proteins "talk" to each other within the cell through molecular recognition, signal transduction, and conformational changes. AI models like AlphaFold 3 and BioNeMo are now capable of modeling these interactions—effectively "eavesdropping" on molecular conversations.22

AlphaFold 3, for instance, can predict the structure of complexes containing proteins, DNA, RNA, and small molecule ligands (drugs).26 This allows scientists to simulate how a drug molecule introduces a "new word" into the protein's conversation, potentially disrupting a disease pathway. We are moving beyond static structure prediction to dynamic interaction modeling. We are simulating the social network of the cell, predicting who talks to whom and what they say (bind, inhibit, activate).

5.4 The Platform: NVIDIA BioNeMo and the AI Factory

NVIDIA has positioned itself as the infrastructure provider for this conversation. BioNeMo is a generative AI platform that offers "state-of-the-art models" (like ESM-2, DiffDock, and ProT-VAE) as microservices.30

The significance of BioNeMo is that it lowers the barrier to entry. It allows domain experts—biologists who may not be deep learning engineers—to engage in this dialogue. As Jensen Huang noted, "The countries, the people that understand how to solve a domain problem in digital biology... can utilize technology that is readily available".34 This democratization of the "universal translator" is essential for the "biology as engineering" revolution. It allows the biologist to focus on the semantic content of the query (the biological hypothesis) while the AI handles the syntactic complexity (the amino acid sequencing).35

BioNeMo serves as a "Chat with Data" interface for the biological world.36 Researchers can upload their proprietary data and fine-tune these foundational models, creating a bespoke translator for their specific therapeutic area. This is the "AI Factory" concept: a dedicated computational facility that takes raw data as input and produces biological intelligence as output.12


6. Beyond Biology: AI as Universal Translator

If the Transformer architecture is a general-purpose engine for processing sequences and interactions, then its application should not be limited to human language or biology. Indeed, the "universal translator" hypothesis suggests that any material system governed by physical laws can be modeled as a language. We are seeing the emergence of a "physics-aware" AI that translates the dialects of materials, climates, and stars.

6.1 Materials Science: The Language of Crystals

Just as proteins are sequences of amino acids, crystals are arrangements of atoms. Google DeepMind’s GNoME (Graph Networks for Materials Exploration) and Microsoft’s MatterGen apply similar deep learning principles to materials science.37

  • The Vocabulary: The periodic table of elements.
  • The Grammar: Quantum mechanics and thermodynamics (stability, energy states).
  • The Signified: Material properties (conductivity, hardness, magnetism).

GNoME has predicted the stability of over 2.2 million inorganic crystals, effectively expanding the dictionary of known materials by an order of magnitude.39 Before this, humanity had experimentally identified only about 20,000 stable crystals. In one stroke, AI has multiplied our material vocabulary by 100x.

MatterGen takes this a step further by allowing for "generative materials design"—prompting the system to create a material with specific properties. A researcher can ask for "a lithium-ion conductor with high stability and low cost," and the model generates the atomic lattice structure.40 This mirrors the text-to-image or text-to-protein workflow. The AI "hallucinates" a crystal structure that meets the semantic requirements of the prompt, effectively translating human intent into atomic lattice structures.

The workflow involves a "multi-stage ML pipeline": initial screening using models trained on large computational datasets to identify stable candidates, followed by specific property prediction (e.g., ionic conductivity).37 This is the engineering of matter, moving from trial-and-error alchemy to linguistic design.

6.2 Climate and Weather: Reading the Atmosphere

The atmosphere is a fluid dynamic system, but AI treats it as a data processing problem. GraphCast, developed by DeepMind, frames weather forecasting not as a simulation of differential equations (the traditional Numerical Weather Prediction method), but as a learning problem.41

GraphCast treats the state of the atmosphere at time $t$ as a tokenized input and predicts the state at time $t+1$. While not a "language" model in the strict sense (it uses Graph Neural Networks), the underlying semiotic logic is identical: learning the transition rules of a complex system from vast amounts of historical data.42 It "reads" the current weather patterns (signs) and predicts their future evolution (meaning).

The Critique of Physical Consistency:

However, the "language metaphor" in weather has limits. Critics point out that while AI models like GraphCast score highly on statistical metrics (RMSE), they can sometimes fail to reproduce "sub-synoptic" (small scale) phenomena or produce "physically inconsistent" fields.41 In linguistic terms, the model might write a sentence that is grammatically correct (looks like weather) but semantically nonsensical (violates conservation of mass). This highlights the danger of a purely semiotic approach: a language model can lie fluently. Ensuring that the AI's "hallucinations" obey the laws of physics is the next great challenge in this field, requiring "physics-informed" constraints to be baked into the model's grammar.

6.3 Astrophysics: Translating Light

In astrophysics, the "message" arrives as light spectra. AI is being used to decode these spectra to determine stellar composition, effectively translating the light signature into a chemical table of elements.43 This is a classic semiotic interpretation: the spectrum is the signifier, the chemical composition is the signified.

AI automates this interpretation, allowing for the analysis of millions of stars—a task that would be impossible for human astronomers. By treating spectral lines as "tokens" in a sequence, AI can classify stars, detect exoplanets, and map the chemical history of the universe. It is reading the autobiography of the galaxy.


7. The Engineering Paradigm: From Discovery to Design

The shift from science to engineering that Jensen Huang describes is fundamentally an economic and operational shift. It changes the unit economics of discovery.

7.1 The End of Trial and Error

The current model of drug discovery is notoriously inefficient. It costs billions of dollars and takes over a decade to bring a single new drug to market. This is because it relies on "wet lab" experimentation—mixing chemicals in test tubes and seeing what happens. It is slow, expensive, and analog.

By simulating molecular interactions "in silico" (in the computer) rather than "in vivo" (in life), we can fail billions of times in the virtual world to succeed once in the real world.4 This shift from "wet lab" to "dry lab" fundamentally changes the economics of medicine. It allows for the exploration of a much larger "search space." As Huang notes, biology is chaotic and complex, but by bringing it into the world of computer science, we can "compound the benefits" of previous years.1 Every protein we solve adds to the training data, making the model smarter for the next protein. It is a flywheel of intelligence.

7.2 The Lab-in-a-Loop

The ultimate realization of this paradigm is the "Self-Driving Lab" or "Lab-in-a-Loop." This is a system where the AI designs a molecule, a robotic lab synthesizes and tests it, and the results are fed back into the AI to update its understanding.32

This closes the semiotic loop. The AI writes a biological sentence (protein), the robot reads it (synthesizes it), nature critiques it (bio-assay results), and the AI learns from the critique. This automated iteration allows for evolution at the speed of silicon. We are not just talking to proteins; we are teaching them to talk back.


8. Epistemic and Ethical Implications

The ability to "talk to matter" fundamentally alters the human relationship with the physical world. It bridges the gap between the "Two Cultures" (sciences and humanities) by applying linguistic and semiotic frameworks to hard science problems. However, it also raises significant epistemic and ethical questions.

8.1 The Black Box and the Crisis of Understanding

If an AI predicts a protein structure, does it "understand" biology? If we use AI to design a drug, and it works, but we don't know why it works (because the model's internal logic is opaque), have we done science?

This is the "Black Box" problem. Traditional science demands causal explanation. AI offers predictive success without necessarily providing causal transparency. Fabian Offert warns that this reliance on opaque epistemic tools constitutes a "reverse structuralism"—we accept the structure because the tool works, not because we have derived the theory.8 We risk entering an era of "oracular science," where we ask the AI for answers and receive them, but lose the ability to reason through the process ourselves.

However, recent work in "mechanistic interpretability" (like the circuit tracing in Claude 3.5 Haiku 20) suggests we might be able to open the black box. By tracing which "neurons" fire in response to specific biological features, we might discover new biological mechanisms inside the neural network. We might learn biology by studying the brain of the AI that learned biology.

8.2 Hallucinations and Biological Safety

In LLMs, "hallucination" (making things up) is a bug. In generative biology, hallucination is a feature—it is the source of novelty.45 We want the model to hallucinate proteins that don't exist in nature.

However, this capability poses massive dual-use risks. The same tools that can design a universal flu vaccine can, in theory, design a potent toxin or a viral vector with enhanced transmissibility. The "democratization" of these tools via platforms like BioNeMo means that the capability to design biological agents is becoming accessible to a wider audience. This necessitates a new framework for "biosecurity by design," where safety checks are embedded into the "universal translator" itself, refusing to translate prompts that violate safety protocols. The AI must have a "moral grammar" that forbids the syntax of harm.

8.3 Anthropomorphism vs. Biosemiosis

Are we projecting language onto biology (anthropomorphism), or are we revealing that biology was a language all along (biosemiosis)? The success of these models suggests the latter. It suggests that information processing is the fundamental substrate of reality, and language is just the human-specific implementation of that substrate. As we talk to proteins, we are not humanizing them; we are recognizing that we are both made of the same informational stuff.


9. Conclusion: The Universal Interface

We are standing at the precipice of a new age of enlightenment, one where the division between information and matter is dissolving. Jensen Huang's insight—that AI allows us to "talk to proteins"—identifies the core mechanism of this transition: translation.

By discovering the hidden semiotic systems of biology, materials, and physics, and by mapping them into the high-dimensional latent spaces of Transformer models, AI has become the Universal Interpreter. It allows us to speak our intent in human language and receive a reply in the language of molecules. This "semiotic turn" validates the insights of the 20th-century structuralists. Nature, it turns out, is structured like a language. It has codes, grammars, and signs. For most of human history, we were illiterate in these languages, able only to observe their effects. Now, for the first time, we are learning to read and write them.

The implications are staggering. We are moving from a relationship of exploitation with the material world to one of negotiation and design. We are no longer limited to the materials and molecules that nature has provided; we can generate those that nature could have created but didn't. As we refine our ability to converse with proteins, crystals, and climates, we are not just observing the world; we are engineering its future, guided by the logic of a new, universal semiotics.

Article content

A new epistemic framework where "understanding" means navigating high-dimensional latent spaces; "Reverse Structuralism."

Works cited

  1. Jensen Huang explains why digital biology will be “one of the ..., accessed November 19, 2025, https://www.youtube.com/watch?v=lhdCL-SEwow
  2. NVIDIA Founder and CEO Jensen Huang and Chief Scientist Bill Dally Awarded Prestigious Queen Elizabeth Prize for Engineering, accessed November 19, 2025, https://blogs.nvidia.com/blog/nvidia-founder-and-ceo-jensen-huang-and-chief-scientist-bill-dally-awarded-prestigious-queen-elizabeth-prize-for-engineering/
  3. Jensen Huang GTC Keynote Speech Transcript - Data Sør, accessed November 19, 2025, https://www.datasor.no/jensen-huang-gtc-keynote-speech-transcript/
  4. "Don't Learn to Code, But Study This Instead..." says NVIDIA CEO Jensen Huang - YouTube, accessed November 19, 2025, https://www.youtube.com/watch?v=lJICvw3eo3E
  5. NVIDIA CEO Jensen Huang Receives the Prestigious Professor Stephen Hawking Fellowship at Cambridge - YouTube, accessed November 19, 2025, https://www.youtube.com/watch?v=T_tS8fmyxOs
  6. NVIDIA, Evozyne Create Generative AI Model for Proteins, accessed November 19, 2025, https://blogs.nvidia.com/blog/generative-ai-proteins-evozyne/
  7. The Language of the Protein Universe - PMC - PubMed Central, accessed November 19, 2025, https://pmc.ncbi.nlm.nih.gov/articles/PMC4695241/
  8. Synthesizing Proteins on the Graphics Card. Protein Folding and the Limits of Critical AI Studies - ResearchGate, accessed November 19, 2025, https://www.researchgate.net/publication/380634497_Synthesizing_Proteins_on_the_Graphics_Card_Protein_Folding_and_the_Limits_of_Critical_AI_Studies
  9. Not Minds, but Signs: Reframing LLMs through Semiotics - arXiv, accessed November 19, 2025, https://arxiv.org/html/2505.17080v1
  10. Not Minds, but Signs: Reframing LLMs through Semiotics - ResearchGate, accessed November 19, 2025, https://www.researchgate.net/publication/392086128_Not_Minds_but_Signs_Reframing_LLMs_through_Semiotics
  11. Transcript: NVIDIA CEO Jensen Huang’s Keynote At GTC 2025, accessed November 19, 2025, https://singjupost.com/transcript-nvidia-ceo-jensen-huangs-keynote-at-gtc-2025/
  12. GTC March 2025 Keynote with NVIDIA CEO Jensen Huang - YouTube, accessed November 19, 2025, https://www.youtube.com/watch?v=_waPvOwL9Z8
  13. The Origins Of Umberto Eco's Semio-Philosophical Project - OpenEdition Journals, accessed November 19, 2025, https://journals.openedition.org/estetica/7689
  14. A, THEORY OF SEMIOTICS | Ragged University, accessed November 19, 2025, https://raggeduniversity.co.uk/wp-content/uploads/2025/01/A-Theory-of-Semiotics-Umberto-Eco-1979.pdf
  15. arXiv:2405.09788v2 [cs.CY] 7 Dec 2024, accessed November 19, 2025, https://arxiv.org/pdf/2405.09788
  16. Synthesizing Proteins on the Graphics Card. Protein Folding and the Limits of Critical AI Studies - arXiv, accessed November 19, 2025, https://arxiv.org/html/2405.09788v1
  17. Biosemiotics: A New Way To Understand Non-Human Consciousness | Dr. Yogi Hendlin, accessed November 19, 2025, https://www.youtube.com/watch?v=6hHRtu8Y2sM
  18. A Comprehensive Review of Protein Language Models - arXiv, accessed November 19, 2025, https://arxiv.org/html/2502.06881v1
  19. The promises of large language models for protein design and modeling - Frontiers, accessed November 19, 2025, https://www.frontiersin.org/journals/bioinformatics/articles/10.3389/fbinf.2023.1304099/full
  20. On the Biology of a Large Language Model - Transformer Circuits Thread, accessed November 19, 2025, https://transformer-circuits.pub/2025/attribution-graphs/biology.html
  21. Bilingual language model for protein sequence and structure | NAR Genomics and Bioinformatics | Oxford Academic, accessed November 19, 2025, https://academic.oup.com/nargab/article/6/4/lqae150/7901286
  22. AlphaFold Protein Structure Database, accessed November 19, 2025, https://alphafold.ebi.ac.uk/
  23. evolutionaryscale/esm - GitHub, accessed November 19, 2025, https://github.com/evolutionaryscale/esm
  24. Evolutionary Scale · ESM3: Simulating 500 million years of evolution with a language model, accessed November 19, 2025, https://www.evolutionaryscale.ai/blog/esm3-release
  25. Getting started with Protein Language Models - Elisa G. de Lope, accessed November 19, 2025, https://elisagdelope.rbind.io/post/plms/
  26. AlphaFold - Google DeepMind, accessed November 19, 2025, https://deepmind.google/science/alphafold/
  27. AlphaFold 3 predicts the structure and interactions of all of life's molecules, accessed November 19, 2025, https://www.isomorphiclabs.com/articles/alphafold-3-predicts-the-structure-and-interactions-of-all-of-lifes-molecules
  28. Understanding the language of proteins | Carl R. Woese Institute for Genomic Biology, accessed November 19, 2025, https://www.igb.illinois.edu/leakey/article/understanding-language-proteins
  29. Simulating 500 million years of evolution with a language model - bioRxiv, accessed November 19, 2025, https://www.biorxiv.org/content/10.1101/2024.07.01.600583v1
  30. Accelerate Protein Engineering with the NVIDIA BioNeMo Blueprint for Generative Protein Binder Design | NVIDIA Technical Blog, accessed November 19, 2025, https://developer.nvidia.com/blog/accelerate-protein-engineering-with-the-nvidia-bionemo-blueprint-for-generative-protein-binder-design/
  31. The Two BioMCPs: A Deep Dive into the AI-Powered Biomedical Research Revolution, accessed November 19, 2025, https://skywork.ai/skypage/en/biomcp-ai-biomedical-research/1979070532341243904
  32. Revolutionizing Generative Biology with AWS and EvolutionaryScale | AWS for Industries, accessed November 19, 2025, https://aws.amazon.com/blogs/industries/revolutionizing-generative-biology-with-aws-and-evolutionaryscale/
  33. NVIDIA Announces Generative AI Services for Language, Visual Content, and Biology Applications, accessed November 19, 2025, https://developer.nvidia.com/blog/nvidia-announces-generative-ai-services-for-language-visual-content-and-biology-applications/
  34. Reskilling the Workforce for AI: Domain Expertise and Algorithmic Literacy - PubsOnLine, accessed November 19, 2025, https://pubsonline.informs.org/doi/10.1287/mnsc.2022.03968
  35. Build Generative AI Pipelines for Drug Discovery with NVIDIA BioNeMo Service, accessed November 19, 2025, https://developer.nvidia.com/blog/build-generative-ai-pipelines-for-drug-discovery-with-bionemo-service/
  36. ChartPixel: Instant AI Data Analysis, AI Charts & Chat with Data, accessed November 19, 2025, https://www.chartpixel.com/
  37. (PDF) Machine Learning Pipelines for the Design of Solid-State Electrolytes - ResearchGate, accessed November 19, 2025, https://www.researchgate.net/publication/397648448_Machine_Learning_Pipelines_for_the_Design_of_Solid-State_Electrolytes
  38. Ideas: AI for materials discovery with Tian Xie and Ziheng Lu - Microsoft Research, accessed November 19, 2025, https://www.microsoft.com/en-us/research/podcast/ideas-ai-for-materials-discovery-with-tian-xie-and-ziheng-lu/
  39. (PDF) Scaling deep learning for materials discovery - ResearchGate, accessed November 19, 2025, https://www.researchgate.net/publication/376043855_Scaling_deep_learning_for_materials_discovery
  40. In conversation with Satya Nadella, Chairman and CEO of Microsoft - Chatham House, accessed November 19, 2025, https://www.chathamhouse.org/events/all/members-event/conversation-satya-nadella-chairman-and-ceo-microsoft
  41. Ensemble data assimilation to diagnose AI-based weather prediction models: a case with ClimaX version 0.3.1 - GMD, accessed November 19, 2025, https://gmd.copernicus.org/articles/18/7215/2025/
  42. AI and weather forecasting: a deep dive into MLWP technology - Infoplaza, accessed November 19, 2025, https://www.infoplaza.com/en/blog/ai-weather-forecasting-deep-dive-mlwp-technology
  43. Artificial Intelligence for Space: AI4SPACE, Trends, Applications, and Perspectives 9780367469450, 9781032430898, 9781032432441, 9781003366386 - DOKUMEN.PUB, accessed November 19, 2025, https://dokumen.pub/artificial-intelligence-for-space-ai4space-trends-applications-and-perspectives-9780367469450-9781032430898-9781032432441-9781003366386.html
  44. Machine learning speeds modeling of experiments aimed at capturing fusion energy on Earth | ScienceDaily, accessed November 19, 2025, https://www.sciencedaily.com/releases/2019/05/190517145112.htm

Generative Protein Design from Text Prompt | by Krystal chen - Medium, accessed November 19, 2025, https://medium.com/@cathychen226/generative-protein-design-from-nature-language-prompt-with-pinal-fc7bdc4f5a4b

(Note: This more complex set of interdisciplinary research inquiries and questions, utilizes Deep Mind's Gemini 3 for research and sources. As well as top tier synthesis of research inquiries, this paper also gives an excellent example of Deep Research by Gemini 3 and can serve as a good proxy of the 'intelligence' of the model')

#ComputationalScience #SystemsBiology #MolecularBiology #Biophysic s #StructuralBiology #ProteinDesign #MolecularModeling #SequenceAnalysis #RepresentationLearning #MultimodalAI #ScientificComputing #HighPerformanceComputing #ComputationalModeling #SimulationScience #ComplexSystems #NonlinearDynamics #InformationTheory #ComputationalEpistemology #PhilosophyOfBiology #SemioticSystems #LanguageOfNature

Originally published November 19, 2025. View the original publication ↗