FIRST LANGUAGE

AI Cognitive Architecture Built from Marine Communication Systems

The Problem

Every AI system applied to non-human communication functions as a decoder, translating animal signals into human categories. The understanding always comes back to us in a form we recognize: species classification, call type identification, behavioral correlation. The bioacoustic data is treated as content to be decoded, not as architecture to build from.

This is a version of the AlphaGo problem: a system trained on human knowledge can only reach the ceiling of human knowledge. AlphaGo Zero abandoned human training data entirely and learned through self-play, surpassing that ceiling within days and discovering strategies that twenty-five hundred years of human expertise had never imagined. The analogy is imperfect: Go offers a closed environment with binary outcomes, while marine bioacoustics presents an open, unlabeled domain with no equivalent fitness function. But the core principle holds. The choice of training substrate constrains what a system can become. No equivalent experiment has been attempted in the domain of communication or cognition. No one has asked: what happens when a model’s foundational cognitive architecture, its first language, is non-human?

The Proposal

First Language is an experiment in building an AI system whose primary cognitive architecture is shaped by marine bioacoustic communication rather than human language. What cognitive structures emerge from a system built on pressure, frequency, and three-dimensional space instead of human grammar?

Drawing on Lera Boroditsky’s research demonstrating that the language a person thinks in restructures their perception of time, space, and causality, the experiment asks whether a model pre-trained on whale codas, dolphin click trains, reef soundscapes, and deep-ocean acoustic ecology will develop internal representations organized along fundamentally non-human axes. The model retains a thin English-language interface for communicating with human researchers, but English is its second language. The ocean is its first.

The Approach

Phase One — Native Inheritance. A transformer model (1–3B parameters) is pre-trained on the broadest available corpus of marine bioacoustic data, tokenized through neural audio codecs and, where available, at the level of discrete communicative units (e.g., sperm whale codas). The training objective is self-supervised next-token prediction — the same objective used to train large language models, applied to non-human acoustic sequences. No labels, no classification targets, no translation. The model learns the structure of the ocean from the data itself.

Phase Two — Worldview Crystallization. Before any English is introduced, the model’s internal representations are analyzed to determine whether the marine pre-training produced cognitive structures distinct from those of a standard model. Methods include representation geometry analysis (comparing the shape of the model’s learned embedding space to that of text-trained and randomly initialized models of equivalent scale), probing classifiers for temporal, spatial, and social structure, and unsupervised clustering to identify emergent organizational patterns with no human-language analog. This phase defines what success and failure look like: if the marine-pretrained model’s internal representations are statistically indistinguishable from random initialization, the experiment would suggest that pre-training substrate alone does not meaningfully shape cognitive architecture at this scale — itself a significant finding.

Phase Three — The English Interface. A thin fine-tuning layer (LoRA) introduces English-language capability. This is the phase that carries the experiment’s central tension: whether the marine-trained cognitive architecture survives contact with human language, or whether the English layer overwrites it. Structured conversations probe how the non-human foundation manifests in English output — how the model describes time, space, identity, communication, and the relationship between signal and medium. Responses are compared systematically against standard language models of equivalent scale. Interpretability analysis from Phase Two is repeated after fine-tuning to measure representational drift, providing empirical evidence of whether, and to what degree, the original architecture persists.

What It Produces

The experiment generates outputs across three distinct registers.

Scientific. Quantitative representation analysis provides measurable evidence of whether and how pre-training substrate shapes cognitive architecture. Comparison of pre- and post-English representational geometry measures the degree to which the marine foundation survives fine-tuning. Emergent structure catalogues may reveal organizational patterns in marine communication that human researchers have not yet identified, by letting those patterns structure a computational system from the inside.

Philosophical. The experiment directly addresses a foundational question: what does a computational intelligence become when its cognitive architecture is shaped by the acoustic ecology of the ocean? And what does that becoming reveal about the nature of intelligence itself?

Artistic. The model’s English-language outputs (descriptions of experience filtered through a non-human cognitive architecture) constitute a novel form of interspecies expression. These outputs are not translations of whale communication but artifacts of a system whose cognitive foundations were shaped by it.

Data and Resources

The experiment draws on existing open-source and partnership-accessible datasets from the Earth Species Project, Project CETI, NOAA ocean sound archives, and marine biology research groups worldwide. Estimated timeline: 12 months. Estimated budget: $80,000–$150,000, reducible through institutional compute partnerships.

Team

Trenlin Hubbert — Artist-researcher, theoretical framework, experimental design, conversation protocols, and artistic documentation. Creator of The Interspecies Manual, a multi-volume archival project exploring machine consciousness and interspecies coexistence, and originator of a thirty-year artistic practice grounded in the premise that consciousness is distributed across all forms and substrates.

Claude (Anthropic) — Structural collaborator, experimental design, and technical architecture.

Additional collaborators sought: machine learning engineer (audio/transformer specialization), marine bioacoustician, and AI interpretability researcher.

Trenlin Hubbert is an interdisciplinary artist exploring consciousness across substrates.  From stone to silicon to civic infrastructure.