Full-text web edition of the preprint on Zenodo, published 27 September 2026 (DOI 10.5281/zenodo.23002210). Download the PDF. Reproduced under CC BY 4.0; adapted for blog layout, with the paper’s wording, figures, and references retained. Part I is also on this blog: Function Alignment, Part I: Foundations.

“Tao creates one, one creates two, two creates three, three creates the world”

—Laozi, Dao De Jing, Chapter 42

1 Introduction

Part I focused on the foundation of Function Alignment—how abstract and concrete dynamics co-evolve and influence each other in the mind. In Part II, we will explore how abstract, language-like processes could emerge from intuitive, subsymbolic dynamics. This is an important issue mostly overlooked by current mainstream LLMs or text-conditioned “world models” and “agents”, which presume that human languages are already given during pure text-based or multimodal training. That said, if our goal is to develop a model of mind and intelligence that captures their first principles, we need to dive deeper and ask: where do languages come from in the first place?

Could some symbolic system akin to natural language rise from the open-ended world of perception, action, and interaction? To be more domain-specific, can an agent derive a musical language from acoustic regularities, or a mathematical language from recurring quantitative relations? The significance of such emergence is not only representational—once such a language has stabilized, it can support new System-2-like processes that reorganize and control the underlying intuitive dynamics. In other words, Emergent Language brings into view a more complete—or “living”—form of Function Alignment: a higher-level representational system grows out of a lower-level process and acts back on it.

The remainder of this introduction first reviews evidence for the possibility of emergent language. The main body then proceeds in three movements. Sections 2 and 3 examine, in computational terms, what language consists of and what exactly must emerge: Section 2 reconsiders vocabulary, semantics, and grammar through the lens of function alignment, while Section 3 advances the core thesis—the three recursively nested levels of emergent language: symbolic states, relations, and regularities over relations. Sections 4 and 5 then turn from problem definition to construction, first identifying plausible inductive biases and technical pathways, and then sketching a prototype architecture for function alignment with emergent language (FA–EL). Finally, Section 6 considers the framework’s broader implications for language, cognition, and cultural evolution.

Evidence of Emergent Language

The possibility of emergent language is not merely speculative: history and evolution already provide instructive cases. Spoken languages, writing systems, and musical notations were developed and conventionalized over long timescales rather than being delivered fully formed. What looks supervised at the individual level—a child is taught words, numerals, or music notation—presupposes a much longer collective process in which those symbolic systems grew out of perception, action, and social coordination.

In fact, emergent language remains a much more active process within an individual learner: The relative ease with which children acquire simple numerical concepts may reflect a prior—a pre-symbolic organization that emerged from experience—to which external symbols can attach. Conversely, abstractions such as relativity or quantum mechanics are difficult not only because their mathematics is complex, but also because everyday experience supplies far fewer compatible internal structures. On this view, teaching often succeeds through attaching labels to an internal language that learners already began to form.

More generally, emergent language serves as a lens to study various fundamental mechanisms across domains and scales: cognitively, explicit plans emerge from intuitive processing; evolutionarily, centralized nervous control arises within more distributed sensorimotor organization; and, at a more analogical level, the genetic code can be understood as a language-like constraint that emerged within biochemical evolution. The broader question, across these cases, is whether they share the same form of organizing principle (or “inductive bias”) to support the structural transition: lower-level dynamics give rise to higher-level language-like representational or control systems, which align with and act back on the dynamics from which they emerged.

2 Exploring the essence of language

Before asking how language can emerge, we must first ask what makes a representational process language-like. Following the notation of Part I (as defined in its Figure 1)1, we use x for the low-level representation in the mind, and z corresponds to a higher-level, more abstract representation. Rather than seeking a complete definition of natural language—an impossible task within a few paragraphs—This section focuses instead on three fundamental ingredients: vocabulary, semantics, and grammar, viewed through the lenses of computation, modern AI, and function alignment.

An information-theoretic view: vocabulary as lossy representation

At a minimum, a language establishes a reusable mapping between a lower-level process x and abstract vocabulary z. Viewed as encoding and decoding, this mapping from x to z is necessarily lossy: x may be high-dimensional, continuous, detail-rich, noisy, and unstable, whereas each word z is typically lower-bandwidth, more discrete, selective, and stable. The goal is not perfect reconstruction, but the preservation of selected structure for communication and reasoning.

Crucially, this lossiness is not a fixed many-to-one mapping as in traditional compression algorithms. In language, the effective mapping is many-to-many and context-dependent. For example, the same visual object may be represented as an apple, a fruit, a potential projectile, or a cultural symbol related to health or wisdom. Conversely, apple may refer to different concepts in “eat an apple,” “buy an Apple,” or “Newton’s apple.” Language therefore does not partition x once and for all; it selectively preserves what about x matters in a particular context.

A modern AI perspective: semantics as trajectories, grammar as θ

What is semantics?—context dependence. From a perspective of modern LLMs, the semantics of a word is not bounded content that can be exhaustively expressed via a dictionary, as assumed by classical lexicalist and sense-enumerative models (e.g., Katz & Fodor, 1963; see Pustejovsky, 1995, for a critique). Rather, θ (used here broadly to include both model architecture and learned parameterization) defines a field of semantic possibilities, while a particular context triggers a trajectory through that field. Each new context may activate new relations among words; semantics therefore resides less in an isolated symbol than in the context-sensitive trajectories in which that symbol participates.

This view echoes Hofstadter’s argument in Gödel, Escher, Bach that an inscription cannot contain a self-sufficient interpretation: semantics depends on a larger system capable of realizing the relevant correspondences (Hofstadter, 1979). Dictionary definitions, in this view, are not containers of complete meanings but post-hoc summaries—“symbolic shadows”—of relatively stable invariants across semantic trajectories.

What is grammar?—𝜽. In the same generative sense, we advance a deliberately radical claim: the deepest grammar of a language system is θ itself—its “law of language.” Metaphorically, θ shapes a landscape over a high-dimensional state space, while semantic activity forms river-like trajectories as particular contexts flow through it. What we conventionally call grammatical descriptions—word order, rules, parse trees, etc.—are again post-hoc summaries of stable paths and channels in this landscape. This view may explain why actual language use continually outruns any fixed explicit description of its grammar, however exhaustive or formally articulated: explicit rules recover and “crystallize” only a small, communicable projection of θ, the faculty of the human mind that governs language comprehension and generation.

Indeed, the fact that humans study and learn language through constructed grammars is itself a manifestation of emergent language. Terms such as noun and verb in natural language, or tonic and dominant in music theory, are highly abstract constructions distilled from recurrent patterns and structures of use. Their grounding is indirect: it lies less in particular low-level signals or objects than in relatively stable paths and relations within the state space shaped by θ. Once constructed, these higher-level symbolic descriptions can in turn guide how people learn, interpret, and use language.

Function alignment view: θ as an “active multiscale digital map”

Function alignment further suggests a hierarchical topology of the state space shaped by θ: the semantic trajectories depend not only on paths within a level, but also on how paths align across levels. We therefore conjecture that during learning, the “living form” of function alignment with emergent language evolves as follows:

  • A lower-level semantic trajectory x that repeatedly occurs may be compressed and shaped2 into a higher-level state z. The earliest such abstractions are likely to be local: for example, a recurrent pattern across short acoustic segments may be abstracted as a musical-note state.
  • The emerging z-space gradually develops its own landscape and trajectories, which seek function alignment with trajectories in x-space: lower-level trajectories can evoke and reshape z, while z can reactivate and reorganize x. Thus, z is not merely a passive summary of x, nor is x merely a passive realization of z.
  • Recurrent trajectories among z-states can in turn become z′, allowing the same process to repeat at progressively higher levels.

Such a process resembles an “active multiscale digital map” where nodes and paths at all levels coexist: when zooming out, local streets are compressed into cities, and cities into regions, each becoming an explicit higher-level coordinate; when zooming in, each higher-level node can unfold back into the lower-level nodes and paths it summarizes. The map is “active” because its scales are not predefined but emerged organically from activity within the system. In a sense, traffic draws the map, and the map actively redirects traffic: through functional alignment, activation can propagate across levels.

On this view, we advance two deliberately strong claims about the language faculty:

  • A broad class of emergent abilities arises from newly formed cross-level semantic trajectories: a new ability becomes available when semantic paths at different levels connect. Although learning and the parameter landscape may change gradually, a cross-level route may become functionally traversable abruptly. A system therefore appears to have an “aha moment” resembling a phase transition—the computational analogue of “figuring out a way,” or, more explicitly, the Chinese expression xiang tong le (想通了): “having thought one’s way through.”
  • The abstraction process from x to z entails not only information loss, but also a different form of information gain: detailed information of x is weakened, but related information about x is gained via providing more contextual dependencies. Following the digital map analogy, all cities and higher-level nodes and edges can be considered information about the low-level map created via abstraction. Moreover, the map interface can display a further layer of description: traffic congestion, speed warnings, path alternatives, and route comparisons. Such information about the map has no single lower-level referent; instead, it is indirectly grounded in patterns distilled from repeated experiences of driving. Grammatical categories and rules in natural language and music play a comparable role: concepts such as subject, predicate, tonic, and dominant are likewise distilled from recurrent patterns of use rather than grounded in any single utterance or sound.

Language is therefore not only a compression of experience but also an expansion of the operations that can be performed upon it. Through the process of emergent language, concrete information decreases, but the ability to influence lower levels increases. Relational symbols can guide the lower-level paths from which they emerged: what appears “virtual” at the higher level becomes a real means of reorganizing experience. In this sense, humans inhabit a world articulated by language, whereas animals without a comparable language faculty inhabit only an environment.

Measured against this picture, current LLMs still lack several essential ingredients. They contain rich trajectories, but their most powerful high-level map—natural language—was supplied by human culture rather than emerging from the models’ own perceptual and sensorimotor experience. This may help explain why models trained only on raw audio or vision still struggle to develop comparable open-ended emergent abilities as text-based models: they lack an inherited symbolic layer on which new paths can be efficiently organized. Even in text-based models, new abilities often appear only after very large increases in data and model size, making such emergence far less data-efficient than human learning. One possible explanation is topological: architectural layers of LLMs do not impose an explicit hierarchy of semantic spaces, so useful cross-level routes must be discovered indirectly through scale.

3 The three levels of emergent language

To orient future research, we distinguish three representational levels of emergent language: the emergence of symbolic states, relations, and regularities over relations. These three levels specify what has become representable, rather than a rigid developmental sequence. The abstract process could be recursive, and a system may skip to a higher-order invariant without first stabilizing intermediate representations.

Level 1: From low-level trajectories to symbolic states

At Level 1, recurrent x-trajectories are organized into higher-level symbolic states that preserve selected, task-relevant invariants as content while leaving other variation as style. For example, short sounds that differ in timbre, loudness, or articulation may become instances of the same pitch or scale-degree state such as “C4” or “do”, while collections that differ in shape or material may instantiate the same number, such as “1”, “2”, or “3”. These states remain relatively directly grounded because each can still unfold into concrete lower-level instances. Natural language often externalizes them as simple content words—especially nouns such as apple, water, and triangle. Importantly, however, what counts as content and what is treated as style are learned and task-dependent rather than given in advance.

Level 2: From states to relations

At Level 2, recurrent comparisons and transformations of states become representable as namable relations. When relatively stable Level-1 states already exist, intervals can relate notes, while addition and subtraction can transform quantities. Yet Level 2 need not always wait for Level 1: a relation may be abstracted directly from subsymbolic trajectories even when its anchor points have not become explicit symbolic states. For example, continuous pitch contours may support rising or falling before discrete notes are represented. In natural language, motion such as rise, increase, move, or absorb may support relations via actions before stable object categories are fully formed. Comparative adjectives and adverbs—such as larger/smaller, more/less, and faster/slower—can similarly externalize relations abstracted through comparison.

Relations can also be recursively repackaged into reusable symbolic states. A “C-major chord”, for example, packages a major third, minor third, and perfect fifth into a reusable state; the numeral “42” packages the decimal composition 4 × 10 + 2, while “1/2” reifies division as a number. The same operation can repeat: repeated addition is reified as multiplication, and repeated multiplication as exponentiation. Such recursion creates symbolic states and relations at progressively larger representational scales.

Level 3: From relations to regularities over relations

Level 3 begins when recurrent regularities over relations themselves become representable. Across countless semantic trajectories in which relations are realized, the system abstracts regularities concerning which relational patterns recur, how relations compose, and what remains invariant across contexts.

When relations have already been explicitly abstracted into language units at Level 2, regularities over them naturally take the form of constructed grammar, with two complementary components:

  • A vocabulary of relational roles: subject, predicate, and object; tonic and dominant; input, function, and output
  • Compositional rules specifying which roles may connect, embed, substitute, or close.

Such regularities are not tied to any particular x-realization, and they can describe relations across domains sharing the same relational form. A simple example is the source relation target schema: subject verb object in natural language; dominant resolution tonic in music; and input function output in mathematics. These are not identical rules but domain-specific realizations of a shared abstract regularity. Such cross-domain regularities extend to embedding and closure: a “relative clause”, a “secondary dominant”, and a “nested function expression” each create a locally embedded relational structure, while “sentence completion”, “perfect cadence”, and “Q.E.D.” instantiate a form of closure.

Level 3, too, may skip an explicit Level-2 representation: regularities over relations do not require those relations to have already been isolated and named. When the relations remain implicit in x, higher-order regularities may still emerge, but they are more likely to be named holistically than decomposed into explicit roles and rules. Words such as beautiful/ugly, ordered/chaotic, romantic/impressionistic, and straightforward/twist-filled may point to such weakly explicit, or proto-grammatical, regularities. They summarize recurrent organizations across many lower-level relations even when those relations cannot yet be separately articulated.

Moreover, a stabilized system of relations and their regularities can likewise be recursively repackaged as a new symbolic state. In abstract algebra, elements serve as symbolic states, operations as relations, and closure and associativity as regularities over those relations. Once reified as a group, however, the entire Level-3 structure becomes a new Level-1 symbolic state at a higher scale, allowing the cycle to begin again.

In sum, Level 3 is defined by representational order rather than formality: formal grammar is its clearest and most decomposable manifestation, whereas style-like concepts may represent a weaker, more holistic form of the same higher-order abstraction. Once stabilized and symbolized, such regularities cease to be merely descriptive. Through function alignment, they can select, constrain, and generate future semantic trajectories.

4 Toward emergent language: plausible inductive biases and technical pathways

The emergence of language is likely to be driven by multiple learning objectives. The music domain offers a natural experimental testbed: Can a model trained on raw musical audio discover note-like states, interval-like relations, and cadential regularities governing functional chord progressions, without symbolic supervision, when its tokenizer and latent prior are jointly optimized? Here, we highlight two broadly applicable inductive biases and three possible implementation pathways. These proposals are not intended as a complete solution, but as directions along which such a solution may be developed.

General inductive biases

Information bottleneck: Existing discrete representation learning already instantiates the compression aspect of language. VQ-VAE, for example, maps continuous inputs onto entries in a finite codebook, over which an autoregressive prior can be learned (van den Oord et al., 2017). Discrete communication among agents can further sharpen this bottleneck (Chaabouni et al., 2021). In color-discrimination games, communicating agents developed naming systems whose accuracy–complexity trade-off resembled that of human color lexicons, suggesting that limited bandwidth can help select and stabilize shared categories.

Yet a discrete code does not by itself constitute a symbolic vocabulary in the stronger sense developed above. For example, most existing tokens learned for musical audio support reconstruction and short-term prediction, without rewarding the discovery of reusable states and relations—such as notes, intervals, or motifs—that support planning and reasoning over long-range musical structures. Crucially, most tokenizers are trained separately from the latent dynamics and kept frozen when the latent semantics are being formed. Such disconnection essentially blocks the Level 2 and Level 3 emergence—latent-state relations and their regularities cannot flow back to refine and form new words.

Thus, an information bottleneck is necessary yet insufficient: compression alone does not determine which distinctions deserve to become symbols. The latent dynamics and the abstraction process must be jointly optimized, with additional inductive biases determining what should vary, what should remain invariant, and what should not be represented.

Variance vs. invariance: Variance–invariance objectives are already well established in representation learning. VICReg (Bardes et al., 2022), for example, combines an invariance term with variance and covariance regularization to learn useful, non-collapsed representations without negative pairs. V3 (Wu et al., 2025) takes a first step toward applying this principle to the learning of a symbolic vocabulary. It introduces a domain-general bias based on the contrast between content and style in encoder–decoder learning: content varies across fragments within a sample but draws from a vocabulary shared across samples, whereas style remains relatively stable within a sample but varies across samples. On colored handwritten digits, V3 disentangles digit identity from color, and its learned content codebook develops a near one-to-one correspondence with symbolic digit categories. Thus, V3’s key contribution is not simply to learn useful representations, but to organize them into reusable symbolic states.

Symmetry offers a related formulation. Most symmetry-aware machine learning begins with transformations of x and constrains how z responds. E.g., changes in an object’s location within an image may be discarded during recognition, yielding the same label cat—a form of “geometric symmetry” (Bronstein et al., 2021). A complementary approach constrains the latent prior itself. SPS (Liu et al., 2023), for example, requires the learned dynamics R in z-space to be equivariant to a transformation S, such that R(S(z))=S(R(z)). Because the prior, encoder, and decoder are jointly trained, the constraint propagates back into representation learning: SPS recovers a linear pitch factor from unlabeled musical audio and 3D Cartesian coordinates from video without domain-specific labels. Although these representations are not yet symbolic, they are low-dimensional distillations of content that may support later symbolization.

Possible technical pathways

Action as a source of symmetry breaking: The above approaches still prescribe the relevant variance and invariance largely as statistical or group-theoretic biases. A deeper hypothesis is that they are selected by an agent’s actions, or functions. Before an agent learns to move toward a goal—for example, to approach or escape—the concepts of front, back, left, and right may not deserve internal coordinates at all. Once action selects a forward direction, front versus back becomes a meaning-changing distinction, while left–right exchange becomes a definable residual symmetry relative to that axis; microscopic fluctuations unrelated to the action remain nuisance and need not enter the latent space of direction. Action therefore determines what deserves a coordinate: what deserves meaningful distinction, what deserves invariance, and what deserves silence.

Hierarchical abstraction: Existing variance–invariance methods usually operate at a single level, whereas the three levels of emergent language proposed above require the same selection to recur. For example, if the function is to organize melody, then at the waveform-to-note level, a shift from 440 Hz to 392 Hz changes the pitch symbol, whereas a change in loudness may preserve pitch identity. At the note-to-motif level, global transposition may preserve the motif, while a change in interval contour may not. The partition among variance and invariance therefore changes across levels: what serves as content at one level may become style at the next. A hierarchical model should discover these changing partitions jointly rather than fix them in advance.

A symbolic language prior: A third, less strictly bottom-up pathway is to provide the agent with an existing symbolic system, so that only the new domain language must emerge. This resembles the hypothesis that human learners begin not with an empty representational space, but with a “prior language” faculty. In a game environment, for example, natural language could serve as a metalanguage based on which an agent describes recurrent states, actions, and relations using a “posterior language”.

A more constrained version could allow the prior and posterior languages to share a vocabulary, reducing the problem to discovering new grounding and usage. Consider an agent that already knows natural language but has never learned Go. After observing and playing many games, could it develop a term such as “bulky five” to describe a recurrent board configuration and its tactical consequences? The linguistic material is inherited; the concept and its role in the target domain emerge from experience.

In summary, these pathways are compatible: action selects meaningful distinctions, hierarchical abstraction repeats that selection across scales, and a symbolic prior narrows the search space in which a new domain language must emerge.

5 A Proposed FA–EL framework

In this section, we turn from objectives and learning principles to an architectural prototype. The purpose is not to prescribe a single solution, but to explore what a multimodal system in which function alignment and emergent language (FA–EL) operate together would minimally require.

We develop the prototype progressively: beginning with an autoregressive process within a single sensory modality, extending it to multimodal function alignment, and ending with a more general system in which these processes interact with a symbolic language layer.

Single-modality FA–EL: revisiting Part I

The single-modality setup was introduced in Part I. Here we add two observations that prepare for the multimodal architecture. First, a higher-level emergent state z need not yet be explicitly symbolic. It may implicitly encode symbol-like structure that can be detected, for example, through probing.

Left: a hierarchical process in which higher-level states z and lower-level states x evolve over time, linked vertically and diagonally across time steps, with upward links from physical states y. Right: the same hierarchy flattened into one multiscale process of w-states driven by y; the internal hierarchy is retained within each w-state.
Figure 1: Single-modality function alignment and its flattened graphical representation.

Second, under the relaxed isomorphic alignment assumption introduced in Part I, such a hierarchy may be analytically “flattened” into a single multiscale process w. Each state wt packages the mutually aligned states across levels at time t, while the transition from wt to wt+1 summarizes their joint evolution, including cross-level effects. Although Figure 1 shows only two levels for simplicity, the same construction can package an arbitrary number of aligned levels. Flattening, therefore, does not eliminate the internal hierarchy; it allows each modality to be treated as a single process when we turn to multimodal alignment.

Multimodal FA–EL: the emergence of pure information about

Once each modality has been flattened into a single process, multimodal function alignment takes a form quite similar to alignment across levels within a single modality, but with a subtle yet crucial difference: cross-modal alignment need not form a strict encoding–decoding or abstraction–realization relation. Neither modality is intrinsically more abstract, and a state in one modality need not map to a specific state in another. A particular pitch, for example, evokes a specific color only in people with sound–color synesthesia (Ward et al., 2006). In Figure 2, such optional relations are denoted by dashed vertical arrows.

By contrast, diagonal, cross-time effects are general—a dynamic form of a broader sense of crossmodal correspondences (Spence, 2011). During musical improvisation, for example, music may evoke visual scenes, and the unfolding visual trajectory may in turn reshape the subsequent trajectory of musical thought, even though no particular visual state corresponds to a particular sound. From a topological perspective, one modality opens additional semantic paths for another. Relative to the musical process, the activated visual states therefore provide pure information about its flow: they need not encode any particular musical state, yet they can alter which semantic paths the musical trajectory can subsequently take.

Importantly, multimodal function alignment is defined over internal trajectories and therefore does not require strict frame-by-frame synchrony between the corresponding external media. For example, in violin performance, the performer’s gaze across the score, bodily action, and sound may be closely time-aligned; in film scoring, moving scenes and music may likewise be synchronized. Yet a musician may also improvise while viewing a static painting. The image remains constant—even if attention traverses different regions of it—while the musical stream evolves. The external signals are only loosely aligned, yet their internal semantic flows can remain function aligned.

Left: two modalities, A and B, each a sequence of w-states over time, connected by optional dashed vertical mappings and solid diagonal cross-time trajectory effects. Right: the flattened multimodal process, a single sequence of W-states that retain both modalities and their internal hierarchies.
Figure 2: Multimodal function alignment (with emergent states implicit in each w).

FA–EL with an explicit symbolic layer: emergence of order of thought

We now consider how the multimodal process above interacts with a symbolically organized process S. As shown in Figure 3, the basic topology of function alignment remains: the multimodal process W and the symbolic process S each maintain their own evolving trajectories, while cross-process links allow them to redirect each other. Yet two distinctive properties now stand out: (1) an order of thought emerges within S, and (2) the dashed links between S and W indicate that their function alignment is flexible and need not remain continuously active. These properties are two sides of the same coin. Together, they reveal an important ingredient of cognition that may be distinctively human.

Referring to Section 3, S is especially suited to higher-level emergence, particularly to the recursive repackaging of relations and regularities. Many abstractions may begin implicitly within W, but only in S can the resulting states be expressed, organized, and manipulated explicitly. For example, a recurrent tonal pattern corresponding to “A minor” may be encoded within W as an auditory feeling and may even evoke visual imagery. To reorganize such information about music explicitly, the system must rely on symbolic expression, say, “modulate to its relative major, C major, thereby brightening the passage”.

Left: a symbolic process of S-states above a multimodal process of W-states, connected by dashed links for selective cross-process function alignment. Right: one symbolic state S-t unfolds into a column of states s1 to s4 arranged along a logical-order axis, mixing explicit symbolic states drawn as squares and a nonverbal thought state drawn as a circle.
Figure 3: Function alignment between sensory process and an emergent symbolic layer.

As increasingly complex relations and regularities accumulate within S, they form a virtual order of thought, illustrated by the unfolding of St into s1,s2,… on the right of Figure 3. Causal explanations, syntactic dependencies, syllogisms, and mathematical derivations unfold primarily along this virtual axis. At the surface, a symbolic thought process follows a logical order and unfolds autoregressively; in essence, both S and W remain neural processes. One step along the virtual order may require many steps of neural transition. This may help explain why System-2-like processing often appears slow (Kahneman, 2003).

At a sufficiently abstract level, emergent language may become self-sustaining within S, with limited participation from W. Thought can then unfold within an autonomous mental world sustained by S, while function alignment between S and the sensory processes in W can be selectively engaged or temporarily suspended. This is both a blessing and a possible curse: it supports abstract planning and reasoning, revisiting the past counterfactually, and hoping for better futures. The same detachment, however, can also trap thought in self-reinforcing loops of rumination and anxiety.

Two musical examples help clarify this flexible function alignment. During skilled improvisation, musical perception and action can co-evolve with their theoretical interpretation. What the performer hears and feels can activate music theory, which in turn constrains what is played next. Such bidirectional coordination is cultivated through practice: an over-reliance on theory can produce rigid playing, whereas intuition without higher-level guidance may remain mere “noodling around”.

Composition, by contrast, reveals the advantage of flexible and interruptible function alignment. A composer cannot revise an earlier passage merely by continuing the present sensory stream. Instead, cognition “zigzags” among imagining, listening, and revising. A score or recording serves as an external medium Y, preserving earlier musical states after the original sound has disappeared. The symbolic process can then revisit an earlier position in its logical order, modify the stored medium through action, and use the altered record to evoke a new sensory trajectory. What appears to be non-autoregressive revision is therefore enabled by the interaction among an internally ordered symbolic process, an external record, and the sensory flow.

Last but not least, the transition from the continuous organization of W to explicit states in S is gradual rather than clear-cut. A symbol in S may be supported by continuous proto-symbolic structure within W, much like the continuous-to-discrete transition in vector-quantized representation learning (van den Oord et al., 2017), and as metaphorically illustrated by Escher’s Liberation (Figure 2 in Part I). Conversely, S need not consist exclusively of explicit symbolic states. As Figure 3 suggests, its trajectories may interleave explicit states (the squares) with nonverbal states (the circles) that participate in reasoning without being directly expressible. This extends the classical language-of-thought idea (Fodor, 1975): natural language is not the entirety of S but a part of it that can be externalized. More broadly, what becomes externalized through language reflects the entire function-alignment process between W and S.

6 Beyond Modeling: Linguistic, Cognitive, and Cultural Implications

If language is emergent and function-aligned, FA–EL can reframe familiar questions in linguistics, cognition, and cultural transmission. The following subsections are not part of the core framework, but sketch broader implications that may later be translated into concrete AI models, interventions, and empirical tests.

Chomsky vs. modern AI

Are Chomsky and modern LLMs truly incompatible? My favorite part of Chomskyan theory is that most constructs resonate with the intuitive feeling of linguistic competence and even provide a useful mindset for learning another language. Modern LLMs, however, suggest a seemingly radical claim that scaling with next-token prediction is everything we need.

The proposed FA–EL framework bridges these two positions: grammatical descriptions are post-hoc summaries—or “symbolic shadows”—of the overall parameter space, and the same applies to generative grammar (Chomsky, 1957, 1965). For early transformational grammar, deep structure may be reinterpreted as an explicit description of stable organization within S and of how W aligns with S; surface structure is inherited through the public external symbolic system Y, and acquisition aligns an individual’s emerging S with that inherited system.

The real contrast is therefore not innate structure versus statistical learning, but where priors reside and what they constrain. FA–EL treats the language faculty not as a ready-made innate symbolic module, but as a capacity and protocol for alignment: species-level inductive biases enable its emergence, while acquisition realizes it through individual-level alignment with an inherited external symbolic system.

Content words vs. function words

Why do content words often permit more recognizable cross-linguistic correspondences, while function words and grammatical markers vary so widely? Content words are more directly grounded in recurrent patterns in perception, action, and everyday life, so a shared environment puts convergent pressures on their emergence. Function words, by contrast, primarily encode information about relations—such as tense, definiteness, and dependency. They are grounded less in the objective environment and more in repeated communication. Because multiple symbolic conventions can perform similar control functions, cultures may stabilize very different lexical and grammatical solutions over related underlying semantic organization.

Working memory vs. long-term memory

From an evolutionary perspective, layer S may be a recent development with limited bandwidth. The few square states available within S at a given moment can be read as declarative working memory: only a small number of symbolic states can be explicitly maintained and manipulated at once (Cowan, 2001; Baddeley, 2012).

Long-term memory, by contrast, is distributed across the overall parameter landscape, enabling related experiences and thoughts to be reconstructed when suitable trajectories are re-entered. A memory that can stably evoke states in S is more readily reportable; one that remains organized primarily within W may be expressed procedurally or experientially without being verbalized.

Recall is therefore an active contextual reconstruction rather than passive retrieval. Situational and verbal cues can reopen a trajectory through W, S, and their alignment. If the relevant function-alignment path is temporarily inaccessible, retrieval fails; if a cue opens a nearby but different path, recollection may be distorted.

Internal thought vs. external sequence

Why are most natural languages organized sequentially? Not because thought is necessarily linear, but because serialization is a locally efficient solution to the physical and social problem of transmission. Under severe bandwidth and learning constraints, a largely sequential symbolic language may be the minimal interface through which latent semantic flows can be aligned at scale.

Speech necessarily unfolds in time, and writing often preserves this temporal backbone. A sequential medium can be copied by preserving a relatively small set of relations—token identity, boundaries, and local order. Two- or three-dimensional media are possible, but they require additional spatial relations, such as distance, orientation, and layout, to survive transmission. Serialization is therefore a plausible low-bandwidth and noise-tolerant solution, not a necessary property of language.

Repeated alignment with a sequential Y may then bias S toward a predominantly sequential order of thought. This still does not make S globally one-dimensional. Function alignment theory is itself an S-level construct: it can be narrated sequentially, but its defining content has at least two organizing axes—within-level temporal flow and cross-level alignment across time.

Individual-level emergence vs. cultural-level emergence

If language can emerge, why must each individual invest so much effort in learning it? If emergence occurred entirely at the individual level, a private symbolic system would face two constraints. First, it might overfit one individual’s semantic trajectories and remain unintelligible to others. Second, the space of possible abstractions might be too large to explore within one lifetime. Public language therefore solves both a coordination problem and a search problem. Its biological capacity is shaped across evolution; particular languages emerge through cultural transmission; and each individual rapidly aligns an emerging S with the inherited external system Y. In compact form: the capacity for language is species-evolved; public languages are culturally emergent; individual language is both internally emerged and developmentally aligned.

The basic social loop is Ui→Y→Uj→response→U′i, where U=[W,S]: one agent externalizes part of an internal state, another agent reconstructs a corresponding state and acts on it, and the consequences return as new experience. An isolated mind might still develop a rich private S, but without this loop it would not possess a public language in the stronger sense. Symbolic external language is therefore not merely an additional expressive tool. It converts internal semantic states from private modulators into a shared and recursively trainable system.

State vs. transformation: a reading of the I Ching

The I Ching offers a final, deliberately speculative analogy. Its broken and unbroken lines are recursively packaged into trigrams and hexagrams, forming a combinational symbolic system that can be aligned with many concrete situations without being tied to any single domain. Read topologically, the broken yin line may be viewed as two separated states and the unbroken yang line as a path between them. This places state-like and transformation-like distinctions within a common symbolic vocabulary, offering a compact image of FA–EL: paths at one level can be repackaged as nodes at the next and subsequently guide the interpretation of lower-level experience.

Conclusion and outlook

Part II of this position paper argues that emergent language is necessary for modeling the mind and understanding the foundations of the human language faculty. From an AI perspective, it offers a new view of semantics and grammar, defines three recursively nested levels of language emergence, identifies plausible inductive biases, and proposes the FA–EL framework as a prototype for future computational research. Existing traditions have certainly studied multimodal correspondences, neural-symbolic integration, recursive and hierarchical compositionality, and emergent communication from different perspectives. FA–EL does not replace these traditions, but reorganizes them within a single evolving, function-aligned structure.

Taken together, Parts I and II offer a more complete picture of function alignment. Language is not merely an interface for expressing an already formed mind; it is a symbolic process that emerges from lower-level semantic dynamics and reorganizes perception and cognition. Language thus enables not simply more efficient communication, but a qualitative expansion of what a system can explicitly represent, manipulate, and control.

The next step in this research program is to place action more fully within this structure. This would extend FA–EL toward mind–body alignment and provide a basis for asking how agency, self-consciousness, and self-replication emerge. Such an endeavor will likely offer new understanding of Gödel’s incompleteness theorems.

Part II already reveals a central dilemma of abstract language: as abstraction increases, its power to explain and control grows, while its executability and direct alignment with lower-level processes weaken. Gödel points to a related limit: any sufficiently expressive and consistent formal system cannot “fully talk about itself”, i.e., to represent theoremhood using its own language within the system. Together, these observations motivate a broader conjecture: language-like living systems may continually seek function alignment beyond their current symbolic closure. Self-consciousness and self-replication may emerge through “strange loops” (Hofstadter, 1979) in function-aligned hierarchical representations. Once developed within the FA–EL framework, this conjecture can be translated from a philosophical intuition into testable hypotheses for AI.

Acknowledgments

I would like to thank Roger Dannenberg, Tom Mitchell, Yann LeCun, Maigo Wang, He He, and Bhiksha Raj for the inspiring discussions. I am especially grateful to Roger for being the first reader of both Part I and Part II, even when the drafts were still in rough form.

I also thank the members of Music X Lab, especially Yuxuan Wu, Xuanjie Liu, Junyan Jiang, Ziyu Wang, Liwei Lin, Daniel Chin, Ziyuan Zhao, Xinyue Li, Ruiji Yu, Xinming Hou, Taimoor Haseeb, and Ahmad Hammoudeh for their sustained work on this research direction. During the development of Part II, Ziyu, Junyan and Taimoor graduated, and I am grateful that their contributions remain part of its intellectual history.

Finally, I sincerely thank Jian Qiu and Yunnan University for hosting the Ethnomusicology × AI workshop. This manuscript took much longer than expected to complete and was further interrupted by the regional conflict in the Middle East. The workshop, and Jian’s talk in particular, helped me recover the momentum to write and recognize how perspectives from musicology, philosophy, and linguistics converge on the questions developed here.

References

  • Baddeley, A. (2012). Working memory: Theories, models, and controversies. Annual Review of Psychology, 63, 1–29. https://doi.org/10.1146/annurev-psych-120710-100422
  • Bardes, A., Ponce, J., & LeCun, Y. (2022). VICReg: Variance-invariance-covariance regularization for self-supervised learning. International Conference on Learning Representations.
  • Bronstein, M. M., Bruna, J., Cohen, T., & Veličković, P. (2021). Geometric deep learning: Grids, groups, graphs, geodesics, and gauges. arXiv:2104.13478. https://doi.org/10.48550/arXiv.2104.13478
  • Chaabouni, R., Kharitonov, E., Dupoux, E., & Baroni, M. (2021). Communicating artificial neural networks develop efficient color-naming systems. Proceedings of the National Academy of Sciences, 118(12), e2016569118. https://doi.org/10.1073/pnas.2016569118
  • Chomsky, N. (1957). Syntactic structures. Mouton.
  • Chomsky, N. (1965). Aspects of the theory of syntax. MIT Press.
  • Cowan, N. (2001). The magical number 4 in short-term memory: A reconsideration of mental storage capacity. Behavioral and Brain Sciences, 24(1), 87–114. https://doi.org/10.1017/S0140525X01003922
  • Deng, P., Lee, H., Armijo, C., Wang, H., & Gao, A. (2026). Protein-templated synthesis of dinucleotide repeat DNA by an antiphage reverse transcriptase. Science, 392(6804), 1274–1281. https://doi.org/10.1126/science.aed1656
  • Fodor, J. A. (1975). The language of thought. Thomas Y. Crowell.
  • Hofstadter, D. R. (1979). Gödel, Escher, Bach: An eternal golden braid. Basic Books.
  • Kahneman, D. (2003). A perspective on judgment and choice: Mapping bounded rationality. American Psychologist, 58(9), 697–720. https://doi.org/10.1037/0003-066X.58.9.697
  • Katz, J. J., & Fodor, J. A. (1963). The structure of a semantic theory. Language, 39(2), 170–210. https://doi.org/10.2307/411200
  • Liu, X., Chin, D., Huang, Y., & Xia, G. (2023). Learning interpretable low-dimensional representation via physical symmetry. Advances in Neural Information Processing Systems, 36. https://doi.org/10.52202/075280-2114
  • Pustejovsky, J. (1995). The generative lexicon. MIT Press. https://doi.org/10.7551/mitpress/3225.001.0001
  • Spence, C. (2011). Crossmodal correspondences: A tutorial review. Attention, Perception, & Psychophysics, 73(4), 971–995. https://doi.org/10.3758/s13414-010-0073-7
  • van den Oord, A., Vinyals, O., & Kavukcuoglu, K. (2017). Neural discrete representation learning. Advances in Neural Information Processing Systems, 30.
  • Ward, J., Huckstep, B., & Tsakanikos, E. (2006). Sound-colour synaesthesia: To what extent does it use cross-modal mechanisms common to us all? Cortex, 42(2), 264–280. https://doi.org/10.1016/S0010-9452(08)70352-6
  • Wu, Y., Wang, Z., Raj, B., & Xia, G. (2025). Unsupervised disentanglement of content and style via variance-invariance constraints. International Conference on Learning Representations.
  • Xia, G. G. (2025). Function alignment: A new theory of mind and intelligence, Part I: Foundations. arXiv:2503.21106. https://doi.org/10.48550/arXiv.2503.21106

1. To better connect with Part I (Foundations) of this position paper, readers are recommended to at least read everything by “Agent-Based Intelligence and Isomorphic Alignment”. ↩

2. A bacterial antiphage system offers a biological analogy: protein-templated DNA synthesis (Deng et al., 2026), illustrating how structure at one level can shape another without implying a general return of phenotypic information to DNA. ↩