TECHNICAL FOUNDATION · HERITAGE NEXUS INC.

The Cognitive Continuity Problem

How accumulated human judgment disappears at the moment of greatest consequence, and the architecture we built to prevent it.

PUBLISHED MAY 2026REVISED JULY 2026VERSION 3.0HERITAGE NEXUS INC.BASALITH.AI
v3.0 audit: seven mechanism claims corrected against the running system. See the audit note below.
AUDIT NOTE (v3.0)
Version 3.0 follows an audit of this paper against the running system. Every mechanism claim was checked against the code that implements it. Seven claims did not survive. Each is corrected below and each correction is described rather than quietly removed. The claims that did survive are unchanged. The paper described eleven cognitive dimensions, including Cultural Heritage and Defining Philosophy. The system scores ten, and that set does not include those two. The count has been corrected to the ten the system uses, and the architecture diagram no longer claims each training pair is mapped to a dimension. Section 02 described each dimension's score as an accuracy measure reflecting specificity. It is a coverage measure. It rises with how much of the person's material falls in a dimension, not with how accurately the model reproduces them. Section 02 now says that plainly. Section 04 published a five-part scoring rubric of emotional depth, uniqueness, temporal grounding, relationship context, and specificity. The pipeline does not use that rubric. It scores four axes: specificity, authenticity, trainability, and length. Section 04 now lists the four axes the pipeline actually applies. Section 04 also described a two-tier scoring system in which pairs in a middle score band were escalated to a more capable model for secondary review before entering training. That mechanism was never built. The pipeline runs one scoring pass with a single model. Section 04 now describes only that. Section 06 presented a pre-deployment validation baseline: several validation archives, an average pair density per dimension, a voice and photograph corpus, a contributor network, and languages validated. Those figures did not describe real archives. The block has been removed. Section 06 now reports only the current status of the real early archives. The pair counts previously reported in Section 06 were also wrong in both directions. The total understated the real corpus. The count described as meeting the quality bar overstated it, because it was measured against a threshold the pipeline does not use. Both figures now come from the live system. Section 07 described a regression testing protocol in which the owner approves a baseline suite of responses at onboarding and at each milestone, and new models are tested against that baseline before migration, with divergence triggering human review. That protocol was not built. What exists is a fidelity evaluation that holds out a portion of the person's real deposits and uses a judge model to score how faithfully the entity reproduces them. Section 07 now describes that, and only that. Section 05 described a behavioral fine-tuning layer as one of two permanent parallel systems, and Sections 03, 08, and 12 treated it as built architecture. It is not built. No fine-tuning code, job, or pipeline exists, no archive has a trained model, and the production entity runs a base Claude model with a system prompt assembled from the archive. Retrieval is the one built system, and behavior is shaped today through that system prompt, not through trained weights. Those sections have been corrected to describe what runs. Where the paper discusses fine-tuning as a future direction, it now does so only in the roadmap. Two citation corrections also ship in this version. The cognitive-fingerprint figure of 96.5% was stated as if it held across the cognitive dimensions. It belongs only to the random number generation task in the source study, and Section 11 now matches the scoping Section 02 already carried. The succession cost figure was attributed through a certificate landing page rather than a primary source and was described too broadly. It is now cited to Harvard Business Review directly and scoped to the S&P 1500 population the study measured. The paper corrects itself in public. That pattern stays.
SECTION 01

The Problem Nobody Has Solved

Every person who has ever lived carried irreplaceable knowledge inside them. Not information in the sense of facts and dates. Those can be written down, indexed, retrieved. Something harder than that. The way they evaluated risk. The pattern they recognized in a failing relationship before the other person did. The instinct about a person in the first five minutes that proved correct thirty years later.

This is tacit knowledge in the formal sense defined by philosopher Michael Polanyi in 1958: knowledge that cannot be adequately expressed in words, that exists only in the practice of the person who holds it.[1] It is the knowledge that builds companies, holds families together, and shapes the character of the people around its owner. And it disappears completely at the moment of death or cognitive decline. Precisely when it is most needed.

"We can know more than we can tell."

MICHAEL POLANYI · The Tacit Dimension, 1966

The research on this loss is unambiguous. A 2024 systematic review published in Heliyon analyzed 28 studies on organizational knowledge transfer and found that the loss of tacit knowledge during generational change is one of the defining challenges of 21st-century organizations.[2] Harvard Business Review estimated that badly managed CEO and C-suite transitions destroy close to $1 trillion a year in market value among S&P 1500 companies alone. The authors attribute part of that loss to intellectual capital that departs with the executive.[3] That analysis covered large public companies. The businesses Basalith serves are far smaller, and the judgment is concentrated in fewer people.

For families the calculus is not financial but it is no less real. Research presented at the 2025 CHI Conference on Human Factors in Computing Systems found that participants overwhelmingly identified the fading of family memory and the loss of generational wisdom as the central fear driving interest in cognitive preservation technology.[4]

The problem has three dimensions that previous approaches have failed to address simultaneously:

DATA BLOCK: THE THREE FAILURE MODES
REACTIVE NOT PROACTIVE
Existing digital legacy products work from data that was never intended to train a cognitive reference: texts, emails, social media posts. The person never participated in their own preservation. The result is a reconstruction built from noise.
STATIC NOT EVOLVING
Traditional knowledge capture produces documents. Documents do not improve as AI advances. A PDF written in 2020 is no more useful in 2040 than it was when it was written.
BROAD NOT PERSONAL
Generic AI models learn from humanity in aggregate. They can approximate a type of person. They cannot approximate a specific person: their specific judgment, their specific voice, their specific way of weighing competing values.

Basalith was built to address all three failure modes. The primary user participates intentionally. The cognitive reference is built from training data sourced exclusively from one specific person. And the underlying model layer improves continuously as AI advances, without requiring new input from the person after their death.

SCOPE NOTE
This paper addresses both consumer legacy preservation and organizational succession applications. These are distinct markets with shared underlying technology. For enterprise-specific treatment see basalith.ai/succession.
SECTION 02

The Cognitive Fingerprint

The scientific basis for Basalith rests on a well-established phenomenon in cognitive neuroscience: the cognitive fingerprint. Research published in Scientific Reports demonstrated that individual behavioral patterns in controlled domains are measurably distinctive, establishing that cognitive signatures differ meaningfully between individuals.[5]

The study measured consistency in random number generation sequences, a constrained behavioral domain. Basalith applies the underlying principle, that individuals exhibit stable, distinctive behavioral patterns, across ten broader dimensions of human cognition. Measuring consistency across domains as complex as Approach to Money or Relationship to Family presents a significantly harder problem than the original study addressed. We do not claim 96.5% accuracy across these dimensions. We claim that the patterns exist, are stable over time, and are meaningfully capturable through the methods described in Section 4.

Basalith builds an approximation of this fingerprint across ten dimensions of human cognition. The word approximation is deliberate. What Basalith produces is an algorithmic reference model: a structured representation of how a person has been observed to think, decide, and respond. Not a simulation of consciousness. Not a reconstruction of a person. This distinction governs every design decision in the system.

CORRECTION NOTE (v2.0)
Earlier drafts described the entity as "not a simulation." This framing is imprecise. The Basalith entity is an algorithmic approximation of cognitive patterns derived from a structured training dataset. It responds to novel scenarios based on learned patterns. This is a form of simulation in the computational sense, bounded by the quality and scope of the training data. We use "cognitive reference model" throughout this document to reflect the appropriate epistemic humility.
THE TEN ENTITY DIMENSIONS
01
Early Life
Where the person came from and what formed them: origins, childhood environment, the earliest shaping influences.
02
Core Values
What the person believes most deeply about how to live. The principles that govern their decisions.
03
Approach to People
How the person reads others, builds trust, and handles relationships.
04
Professional Philosophy
How the person thinks about work, leadership, and building things.
05
Approach to Money
What money means to the person and how they weigh risk, security, and financial judgment.
06
Relationship to Family
How the person thinks about family, love, and belonging.
07
Fears and Vulnerabilities
What the person is afraid of and where they feel most exposed.
08
Defining Experiences
The moments that shaped who the person became.
09
Wisdom and Lessons
What the person knows now that they wish they had known earlier.
10
Spiritual Beliefs
What the person believes about meaning, purpose, and faith.

Each dimension carries a coverage score from 0 to 100. The score reflects how much material touches that dimension: how many of the person's deposits fall in it, how many entity responses in that dimension the person has confirmed as accurate, and how many labeled photographs reinforce it. It measures how well covered a dimension is. It does not measure how accurately the model reproduces the person.

Human cognition is inherently contradictory. A person's Approach to Money may conflict sharply with their Core Values under financial pressure. The Basalith architecture does not resolve these contradictions. It preserves them. A person who held genuinely conflicting values is imperfectly represented by a model that smooths those conflicts away. The goal is fidelity, not coherence.

"The most revealing training data is not where a person was consistent. It is where they were not, and how they lived with that."

SECTION 03

System Architecture

The Basalith platform is built on a multi-layer architecture designed for long-term data integrity, real-time cognitive reference interaction, and continuous learning from multiple input modalities.

SYSTEM ARCHITECTURE: INPUT / PROCESSING / RESPONSE
LAYER 01 / INPUT (ALL SOURCES CONTINUOUS)
VOICE DEPOSITS
Phone, portal, app recordings. Whisper ASR transcription.
TEXT DEPOSITS
Written prompts, email replies, wisdom session answers.
CONTRIBUTOR NETWORK
Family and colleague observations. Corrects self-report bias.
PHOTOGRAPH LABELS
AI era estimation plus human narrative annotation.
↓↓↓↓
LAYER 02 / PROCESSING
QUALITY SCORING
Four axes: specificity, authenticity, trainability, length, combined into one score. Only pairs clearing the inclusion gate are used for training.
DIMENSION COVERAGE
The system tracks how well each of the ten dimensions is covered and steers new questions toward the weakest ones.
↓↓
LAYER 03 / RESPONSE (RETRIEVAL AND BEHAVIORAL SHAPING)
RAG LAYER (RETRIEVAL)
Retrieval-Augmented Generation. The permanent factual ground truth. Retrieves relevant deposits at query time. Active from day one through the life of the archive.
BEHAVIORAL LAYER (SYSTEM PROMPT)
System prompt assembled per archive from retrieved material. Shapes tone and expression on the base model. Operates alongside RAG, not in replacement of it.
STACK: Vercel (edge compute) / Supabase (PostgreSQL + RLS) / Anthropic API / Private storage buckets

The critical architectural point, addressed in detail in Section 5, is that retrieval and behavioral shaping are not sequential stages. The RAG layer provides factual grounding. The system prompt shapes expression. Both run on the base model, and neither replaces the other.

The stack: Vercel for edge compute, Supabase with PostgreSQL and Row Level Security for data persistence, and the Anthropic API for all language model operations. All storage is in private buckets. All tables enforce RLS at the policy level, independent of application code.

SECTION 04

The Training Pipeline

The central technical challenge of building a cognitive reference model is not data collection. It is data quality. The Basalith training pipeline is built on one principle: specificity over volume. A single training pair that captures a person's response to a specific situation, with named people, real places, genuine consequences, is worth more than fifty training pairs of general opinion.

TRAINING PAIR QUALITY SCORING
SPECIFICITY
Does the response reference specific people, places, dates, or events? Generic statements score low regardless of emotional weight. This axis carries the most weight.
AUTHENTICITY
Does it sound like this person speaking naturally in their own voice, rather than formal, generic, or robotic?
TRAINABILITY
How much does it reveal about how this person reasons and sees the world, beyond stating a fact?
LENGTH
Substance needs room. Very short responses are capped. This axis carries the least weight.

Scoring is performed by Claude Haiku in a single automated pass at the moment each pair is created. Haiku rates the pair on a small set of weighted axes and returns one combined quality score. The scoring system is model-agnostic. The rubric and database schema remain stable across scoring-model upgrades.

We acknowledge a known limitation. Authenticity is one of the axes the model scores, and it is the hardest to judge. Telling a natural personal voice apart from formal or generic phrasing is a subtle call, and a fast, low-cost model is bounded in its capacity to make it.

The pipeline runs one scoring pass with a single model. Every pair is scored once, and only pairs that clear the pipeline's inclusion gate enter training. There is no second scoring pass and no senior-model review.

Multi-perspective training is a core differentiator. The contributor network provides data self-report cannot generate. When a contributor's observation is confirmed by the primary user through the Wisdom Exchange correction mechanism, the resulting training pair carries the highest weight in the system.

SECTION 05

RAG and Behavioral Shaping

CORRECTION NOTE (new in v2.0)
This note concerns a layer that was not built. It is kept as a record of how the design was described. Version 1.0 of this paper implied that fine-tuning at 500 training pairs would supersede the need for retrieval-augmented generation, that the model would have internalized the pattern and no longer need to retrieve the archive to respond accurately. This framing was incorrect and has been revised.

Behavioral shaping today is done through the system prompt. For each conversation the system assembles a prompt from the archive: retrieved deposits, the person's own words, relationships, and context. This shapes tone and expression on a base model. It does not modify the model's weights, and no per-archive weight training runs in the current system.

Factual grounding, the specific memories, the named people, the real events, must come from retrieval. This is what RAG provides.

RAG LAYER · RETRIEVAL
What the person said.
Retrieves specific deposits at query time. Provides factual grounding. Prevents hallucination of specific memories. Active from day one through the life of the archive.
BEHAVIORAL LAYER · SYSTEM PROMPT
How the person said it.
Shapes how the entity expresses itself: tone, framing, linguistic patterns. Delivered today through system-prompt engineering on the base model. Operates on top of RAG, not in replacement of it.
CORRECTION NOTE (v2.1)
This note concerns a layer that was not built. It is kept as a record of how the design was described. Version 2.0 of this paper stated that fine-tuning teaches the model "how this person thinks." This overclaims what parameter-efficient fine-tuning at 500 samples reliably produces. At this training volume, fine-tuning teaches characteristic expression patterns: tone, framing tendencies, linguistic signature, and the characteristic ways a person structures uncertainty. It does not reliably encode deep reasoning architecture or novel decision-making under conditions the training data never covered. We use "characteristic expression patterns" throughout this document where prior versions used "how this person thinks." The RAG layer remains the mechanism for factual and episodic accuracy. The behavioral layer shapes how those facts are expressed. Neither claim is weakened by this correction. The overclaim is.

"RAG grounds every answer in what the person actually said. The system prompt shapes how it is expressed. Neither invents what the person never provided."

A response built on RAG alone, with no behavioral shaping, is accurate but generic in expression. Behavioral shaping without RAG would be stylistically plausible but factually unreliable, generating content the person never said. The two work together. Remove either and the result degrades.

This is why the engagement system, daily sparks, contributor questions, memory games, wisdom exchanges, is not an optional engagement feature. It is core infrastructure investment in RAG quality.

SECTION 06

Accuracy, Measurement, and Contradiction

Measuring the accuracy of a cognitive reference model is inherently imperfect. Ground truth in this domain is not a fact verifiable against an external record. It is a judgment: does this response reflect how this person would actually respond?

Basalith uses coverage and feedback signals, plus one direct measure of fidelity:

COVERAGE, FEEDBACK, AND FIDELITY SIGNALS
DIMENSIONAL COVERAGE DENSITY
Each of the ten dimensions carries a coverage score. It rises with how much material falls in that dimension: the person's deposits, the entity responses they have confirmed as accurate, and labeled photographs. A well-covered dimension is treated as more developed than a sparse one.
OWNER CORRECTION FEEDBACK
When the primary user tells the system "that is not what I would say, here is what I would actually say," the correction becomes a training pair of the highest quality class.
CONTRIBUTOR VALIDATION
When multiple contributors independently describe the same behavioral pattern and the primary user confirms it, the validated pattern receives the highest training weight.
FIDELITY EVALUATION
The direct measure of accuracy is the fidelity evaluation in Section 7. It holds out real deposits and uses a judge model to score how faithfully the entity reproduces the person.
CORRECTION NOTE (v2.1, figures updated in v3.0): STAGED VALUE CURVE
The current archive status reflects early-stage onboarding: 93 training pairs across 4 archives, of which 58 clear the pipeline's inclusion gate. These are early real archives, not a validation cohort. An earlier framing of this paper treated reaching 500 pairs as the primary value delivery point, making the product feel incomplete until then. That framing was wrong and has been corrected. The value curve is staged, not binary: 10 or more quality pairs activates a RAG-grounded entity that retrieves specific memories, named people, and real events rather than generating plausible generic content. This is already a categorical improvement over any generic model. 50 or more quality pairs across multiple dimensions produces a reference model with meaningful coverage of the person's values, relationships, and defining experiences. 200 or more quality pairs enables reliable multi-dimensional cross-referencing: the system can surface how this person's approach to money has historically conflicted with their stated values, for example. Beyond a few hundred quality pairs, coverage keeps deepening and the entity reasons from a richer, more cross-referenced archive. There is no fine-tuning activation step in the current system. The entity improves as the archive grows, continuously, not at a single threshold. The cold start problem is real. The solution is not to obscure it but to ensure the product delivers clear, demonstrable value at every stage of the curve.

On contradiction: The Basalith architecture does not resolve conflicting values. It preserves them as data. A person who held genuinely conflicting values is imperfectly represented by a model that smooths those conflicts away. The goal is fidelity, not coherence.

SECTION 07

Continuity Across Model Generations

The most important architectural decision in Basalith is the separation of training data from the model that processes it. The training pairs accumulated over years of deposits are stored in a structured database, independent of any specific AI model. When the underlying model is superseded, the archive does not revert. It migrates forward.

A precise framing: a more capable foundation model, applied to the same training data, will produce more articulate, more contextually sensitive responses. It will express the person's cognitive patterns with greater fidelity. What it will not do, and cannot do, is generate new cognitive content the person never provided. The model becomes a better instrument. It cannot add to what is in the archive.

"The data is the permanent asset. The model is the instrument. As instruments improve, the data they work with becomes more fully expressed. But the data itself does not change."

Model migration introduces a non-trivial risk that this framing understates. A new foundation model is not a passive lens applied to the same data. It brings different baseline reasoning patterns, systemic biases, moral weights, and cross-lingual handling. Applied to the same training data, a fundamentally different architecture may alter the perceived character of the entity in ways that are difficult to predict.

Basalith measures this risk with a fidelity evaluation run per archive. The system holds out a portion of the person's real deposits so they are never used to ground the entity's answers. It then puts the held-out prompts to the entity and uses a separate judge model to score the responses two ways: whether the entity can be told apart from the real person, and whether its factual claims stay grounded in the archive. The results are recorded for each run.

The judge model is a smaller model than the one generating the responses it scores. Its value is independence, not superior judgment. A judge weaker than the system it evaluates will miss failures a stronger judge would catch.

This evaluation does not eliminate the risk of persona drift across model generations. What it gives Basalith is a repeatable, measurable check on how faithfully a given model reproduces the person, so a change in fidelity can be seen and investigated rather than assumed.

This is why the early years of an archive, while the primary user is alive, are irreplaceable. Every deposit made during this period is a permanent asset that will be expressed more fully by every model generation that follows.

Voice portrait generation operates on the same principle. The voice recordings captured during a person's life are the permanent asset. Voice synthesis technology will produce increasingly accurate results as it advances. The recordings do not improve. The technology that processes them does.

SECTION 08

Post-Mortem Governance

Version 1.0 of this paper did not address who controls the archive after the primary user dies. This omission has been corrected here.

An archive that can be altered, commercialized, or deleted by heirs against the owner's wishes is not a legacy tool. It is a liability. The primary user who builds a Basalith archive over years of intentional deposits has a reasonable expectation that what they built will be preserved in the form they built it.

THE FIVE POST-MORTEM GOVERNANCE COMMITMENTS
COGNITIVE PATTERN IMMUTABILITY
After the primary user's death the existing training dataset is locked. Surviving family members and heirs cannot modify, delete, or add to the training pairs that constitute the cognitive reference model.
CONTINUED CONTRIBUTION (LEGACY TIER)
Contributor network members may continue to add memories and observations under the Legacy tier. These are stored and clearly marked as post-mortem additions. They do not overwrite the primary training dataset.
HEIR ACCESS: READ, NOT WRITE
Heirs designated in the Legacy tier may interact with the entity and access the archive. They cannot modify the training data. This is enforced at the database level, not just the application layer.
DELETION: EXPLICIT OWNER REQUEST ONLY
An archive is never automatically deleted. Permanent deletion requires explicit written request from the archive owner before their death, or from a designated executor with documented authority. Heritage Nexus Inc. holds the archive for 12 months after a deletion request before permanent deletion.
COMMERCIAL USE: PROHIBITED
Archive data is never sold, licensed to third parties, or used to train models for other users. No beneficiary may commercialize the archive without documented pre-death authorization from the owner.
CORRECTION NOTE (v2.1): TWO-LAYER GOVERNANCE
A previous version of this paper did not resolve the tension between Section 08 (cognitive pattern immutability after death) and Section 10 (ongoing context injection for organizational succession). An external review correctly identified this as a structural contradiction. It has been resolved here. The Basalith architecture maintains two distinct data layers with separate governance rules: The cognitive fingerprint layer contains the training pairs accumulated during the owner's lifetime: their voice, their values, their judgment patterns, their defining experiences. This layer is frozen at death. No successor, heir, or administrator can modify it. This is enforced at the database level. The contextual intelligence layer contains current situational data injected by successors or designated administrators after the owner's death: business developments, market conditions, organizational changes. This layer is explicitly mutable, separately stored, clearly labeled as post-mortem context, and successor-controlled. When a successor queries the entity, the response draws on both layers. The cognitive fingerprint layer supplies the reasoning style and values. The contextual intelligence layer supplies current situational grounding. The two layers are architecturally separate. Writing to one does not modify the other. A further note on model forward-migration: the entity runs on a base foundation model with a system prompt assembled from the archive. When that model is superseded, there are no per-archive trained weights to regenerate. The raw training database is the permanent asset. The model that reads it is the instrument. The person's data does not change. The instrument that reads it does.

We expect this to be an evolving area of law and ethics as cognitive legacy technology matures. The framework here represents Heritage Nexus Inc.'s current position, which we will update as the field develops.

SECTION 09

Privacy and Data Architecture

THE FOUR PRIVACY COMMITMENTS
DATA SOVEREIGNTY
Archive owners retain full ownership of their data. Basalith is a custodian, not an owner. Data can be exported in full at any time on request.
NO THIRD-PARTY DATA SHARING
Archive data is never shared with third parties, never used to train models for other users, and never sold. Each archive is isolated at the database level via Row Level Security policies enforced independently of application code.
PRIVATE STORAGE BY DEFAULT
All storage buckets are private. Photographs, voice recordings, and documents are served through a proxy endpoint with time-limited signed URL generation. No public URLs exist for any archive content.
PERMANENT PRESERVATION
An archive is never deleted due to non-payment. Financial disruption moves an archive to Resting status: data preserved indefinitely, features suspended.

At the database level, Row Level Security is enforced on all tables. Even a misconfigured application cannot access data across archive boundaries. The database enforces isolation at the query level independent of the application layer.

Research on LLM fine-tuning and privacy identified unintended memorization of sensitive information as a key risk.[7] The Basalith approach mitigates this by maintaining training data in a structured database rather than embedding it in model weights.

SECTION 10

Application to Organizational Succession

The cognitive continuity problem exists in organizations as acutely as it does in families. A 2024 systematic review in Heliyon found that the methods available for capturing tacit knowledge remain inadequate across 68% of organizations studied.[2]

The tacit knowledge most critical to organizational performance cannot be documented in a process manual. The judgment a founder brings to an acquisition decision. The instinct a senior executive has about a key hire. The pattern recognition that has guided a company through three economic cycles.

"Organizations face a potential knowledge vacuum due to the retirement of the baby boomer generation. Effective knowledge transfer strategies remain elusive in many organizations."

IGOA-IRAOLA & DIEZ · Heliyon, 2024
THREE KEY DIFFERENCES FROM PERSONAL LEGACY APPLICATION
DECISION FRAMEWORK CAPTURE
The founding session is structured around decision-making frameworks rather than life narrative. The Legacy Guide captures how the founder made decisions: the factors they weighted, the signals they trusted, the conditions under which they overrode their initial instinct.
SCENARIO TRAINING LIBRARY
Twenty or more business scenarios are deliberately constructed and trained into the entity. These cover the situations most likely to arise in the successor's first years.
ONGOING CONTEXT INJECTION
The successor portal requires ongoing business context to remain practically useful. Quarterly calibration sessions with the Legacy Guide update the entity with current business developments. Without continued context injection, the entity's responses become increasingly detached from current reality. The quarterly calibration is not an optional service enhancement. It is a technical requirement.

For the full enterprise treatment: basalith.ai/succession

SECTION 11

Research Foundations

Five research areas underpin the Basalith architecture:

RESEARCH FOUNDATIONS
COGNITIVE FINGERPRINTING
Schulz et al. (2021) measured individual consistency in random number generation and distinguished same-person from different-person sequences at 96.5% AUC. That figure belongs to that constrained task. Basalith applies the underlying principle, that individuals show stable and distinctive behavioral patterns, across broader cognitive dimensions. It does not claim the 96.5% figure across those dimensions.
PERSONALIZED LLM FINE-TUNING
Simchon et al. (2023) showed fine-tuned models can predict personalities from interview language. Au et al. (2025) provides a comprehensive taxonomy of per-user fine-tuning approaches. Research on LLM privacy risks highlights the importance of quality-screened training data over raw personal data.
ORAL HISTORY PRESERVATION
Research on LLMs for oral history analysis demonstrated effective semantic annotation across 92,191 sentences from 1,002 interviews. The Basalith engagement system operationalizes oral history methodology at the individual archive level.
THANATECHNOLOGY ETHICS
CHI 2025 research identified intentional participation as the critical factor distinguishing authentic digital legacy from reactive reconstruction.
TACIT KNOWLEDGE TRANSFER
Igoa-Iraola and Diez (2024) identified storytelling and narrative elicitation as among the most effective methods for capturing tacit knowledge.
SECTION 12

The Roadmap

The version of Basalith that exists today is the least capable version that will ever exist. Every advancement in foundation model quality, voice synthesis, and cognitive modeling directly improves the fidelity of every existing archive without any action required from the archive owner.

NEAR TERM · 2026
Voice portrait generation in 8 languages
iOS App Store distribution
Successor Portal for B2B clients
WeChat integration for Chinese-speaking communities
MEDIUM TERM · 2027-2028
First per-archive behavioral fine-tuning
Video portrait generation
Real-time voice conversation via WebRTC
Multimodal training incorporating image understanding
Enterprise API for organizational succession programs
LONG TERM · 2029+
Real-time reasoning models producing novel responses in the owner's voice
Generational inheritance: descendant archives building on ancestor foundations
Clinical baseline application: cognitive fingerprint preservation for early-stage dementia research
Academic partnership for longitudinal cognitive fingerprint research
Post-mortem governance legal framework

The clinical baseline application deserves emphasis. The same architecture that preserves cognitive patterns for legacy purposes can establish a documented cognitive baseline before decline begins. A person who builds a Basalith archive at 60 creates a measurable record of how they thought at peak cognitive function.

"You never truly leave if you leave enough of yourself behind."

BASALITH · 2026
REFERENCES

Cited Research

[1]

Polanyi, M. (1966). The Tacit Dimension. University of Chicago Press.

Foundational text establishing tacit knowledge: knowledge that cannot be fully articulated in words.

[2]

Igoa-Iraola, E., & Diez, F. (2024). Procedures for transferring organizational knowledge during generational change: A systematic review. Heliyon, 10(4).doi:10.1016/j.heliyon.2024.e27092

28-study PRISMA review. 68% of organizations attempt tacit and explicit knowledge transfer. Effective methods remain elusive.

[3]

Fernandez-Araoz, C., Nagel, G., & Green, C. (May-June 2021). The High Cost of Poor Succession Planning. Harvard Business Review.hbr.org/2021/05/the-high-cost-of-poor-succession-planning

Badly managed CEO and C-suite transitions destroy close to $1 trillion a year in market value among S&P 1500 companies. Part of the loss is attributed to intellectual capital that leaves with the executive.

[4]

Lei, Y. et al. (2025). AI Afterlife as Digital Legacy: Perceptions, Expectations, and Concerns. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems.doi:10.1145/3706598.3713933

Participants identified AI-generated agents preserving family memories as central value proposition. Interviews conducted June-August 2024.

[5]

Schulz, M-A., Baier, S., Timmermann, B., Bzdok, D., & Witt, K. (2021). A cognitive fingerprint in human random number generation. Scientific Reports, 11.doi:10.1038/s41598-021-98315-y

Same-author vs. different-author behavioral sequences distinguished at 96.5% AUC from 300 data points. Fingerprint stable over one week.

[6]

Brickman, J., Gupta, M., & Oltmanns, J.R. (2025). Large Language Models for Psychological Assessment: A Comprehensive Overview. Advances in Methods and Practices in Psychological Science.doi:10.1177/25152459251343582

Reviewing Simchon et al. (2023): fine-tuned model predicting personality traits from social media posts.

[7]

Unintended Memorization of Sensitive Information in Fine-Tuned Language Models. (2025). arXiv:2601.17480.

LLMs memorize training samples even when seen once. Quality-screened training pipelines recommended.

[8]

Au, S. et al. (2025). A Survey of Personalized Large Language Models: Progress and Future Directions. arXiv:2502.11528.

Comprehensive taxonomy of PLLM approaches. Per-user PEFT paradigm.

[9]

Large Language Models for Oral History Understanding with Text Classification and Sentiment Analysis. (2025). arXiv:2508.06729.

Effective annotation across 92,191 sentences from 1,002 interviews in the JAIOH oral history collection.