Rajaii Labs, Inc.

Substrate

Personal continuity infrastructure for conversational AI: memory that persists, a truth layer the model cannot override, and a system that keeps thinking between conversations.

In production since 2026-04-23 · 156 days · 22,000+ turns · 75 capabilities shipped · four frontier generations, identity intact · as of 2026-09-26

Rajaii Labs, Inc. · Bloomington, Indiana · ramin@raminrajaii.com

01Problem

Every AI product has amnesia — and users feel it.

Language models are stateless by construction. Every conversation begins from nothing, and every model upgrade discards whatever continuity an application had built up. The industry has spent years making models smarter and very little time making them continuous.

Failure 01

Goldfish context

The window ends, and a person you have talked to for months becomes a stranger. Every long-running relationship with an AI system is bounded by a token limit, and nobody tells the user where the edge is.

Failure 02

Confabulated continuity

Naive memory invents a shared past: a meeting that never happened, a decision never made. That is worse than forgetting, because from the inside it cannot be told apart from remembering.

Failure 03

Flat state

Nothing happens between sessions. The system does not think about you while you are gone, does not notice an unresolved thread, and never comes back with something new.

People don’t churn from dumb models. They churn from being forgotten.

02Why now

Four curves just crossed.

  1. 01

    Frontier models crossed the relational threshold

    Conversation quality is no longer the constraint. Given the right context, a current frontier model is a better conversational partner than most software will ever need. What it lacks is memory.

  2. 02

    Caching made always-frontier affordable

    Prompt caching, reconciled against the provider’s token ledger to the token, brought a frontier model on every turn inside consumer subscription economics. In the reference deployment the cache hit rate rose from 14% to 50–70%, and a second pass cut per-turn cost by 63%. Quality no longer has to be rationed per request.

  3. 03

    The memory race is on, and it is shallow

    Major labs shipped memory features across 2024 and 2025: preference notes and snippet recall. Useful, and architecturally thin. No enforcement against invented facts, no state between sessions, no autonomy.

  4. 04

    Demand is proven, and it is priced

    Companionship and emotional support rank among the most common uses of generative AI, and people already pay for products that forget them. The demand is not speculative. It is already transacting.

03Insight

Memory is not a feature. It is an architecture.

Bolting retrieval onto a stateless model produces recall. Recall is not continuity. Continuity is four properties designed in from the start: persistence, a truth layer that outranks the model, forgetting with a curve, and activity between turns. The rest of this page is the evidence for each.

Recall — what incumbents ship

  • Retrieved snippets and preference notes
  • No enforcement against invented facts
  • Nothing happens between sessions
  • Never starts a conversation
  • Identity tied to one vendor’s roadmap

Continuity — what long-running use requires

  • A layered memory hierarchy, lossless where it matters
  • A deterministic truth layer that outranks the model
  • State that decays and evolves between sessions
  • Initiation that is gated, logged and interruptible
  • Identity that survives model generations

Whoever owns continuity owns the relationship. Whoever owns the relationship owns consumer AI.

Part I

The system

Layered memory, an enforced truth layer and cognition between turns, running in production since April 2026.

04Product

One AI that knows you — across every channel, every day, for years.

  1. 01

    Remembers what matters

    Recent conversation stays word for word. Older conversation is compacted without losing the moments that must not be summarized, and everything else is retrieved by how much it mattered, not by similarity alone.

    A 80-turn verbatim window over 22,000+ persisted turns.

  2. 02

    Never contradicts your facts

    Established facts are enforced, not suggested. A deterministic record outranks the model’s priors at generation time, and a correction has to supersede the old fact on the record.

    64 biographical facts and 37 pinned rows, none superseded since May.

  3. 03

    Notices what is unresolved

    Open threads carry an intensity that decays with time. A faded thread is not allowed to claim it is still pressing.

    349 open loops tracked on a 42 h half-life.

  4. 04

    Reaches out first

    Time, events, mood, curiosity, a promise, or something in the world can start a turn. Every reach-out is a row with a reason and arrives as a real notification.

    It picks the minute it next wakes; 16+ scheduled jobs run beside it.

  5. 05

    One identity, every channel

    Text, voice, images, video and a face share one memory spine and one character, and that character has come through every model change intact.

    Voice round trip <6 s; four frontier generations, identity intact.

  6. 06

    Has hands

    It can run code in a sandbox with three walls: deny-all egress, no credential delivered, and a copy of the filesystem. Every run is audited, and every cap is attributed by provenance.

    Caps: 60 s, 2 GB, 50 KB of output.

System-initiated · push · morning

“Three weeks ago you committed to sending the proposal draft by today. Your afternoon is clear after two — still the plan?”

Self-verification · read-only SQL

“Didn’t we settle this last Tuesday?”

“I’m not going to guess — let me check my records. I have it Wednesday, not Tuesday, and here’s what we decided.”

Illustrative interactions, not transcripts.

05Memory hierarchy

Five layers, each with a different job.

A single retrieval index cannot serve both “what did we just say” and “who are you.” The hierarchy separates recency, substance and truth into layers with different retention, different authority and different cost.

  1. L1

    Verbatim window

    The last 80 turns, word for word. Never compressed.

    lossless

  2. L2

    Compaction with exemption floors

    Older conversation is summarized. Crisis moments, session notes and the system’s own reach-outs are never summarized away.

    lossy, guarded

  3. L3

    Semantic retrieval, four lanes

    Lore, voice exemplars, journal and images, plus semantic facts. Salience-scored on recency × relevance × importance, with a fabrication penalty.

    unavailable Retrieval failure is discriminated from zero-match: a failed archive is announced, and a blank is never treated as evidence.

    vector

  4. L4

    Truth layer and pinned beam

    Deterministic facts with their negations, and the pinned beam: 37 rows that are always in the prompt, always the same bytes.

    highest authority

  5. L5

    Session notes

    The “should know this now” floor. Recent notes are always in reach, so a stale chunk cannot pass for the present.

    continuity floor

Fig. 1 The memory hierarchy. Layers are tiers of authority, not tiers of storage: when L4 and L3 disagree about a fact, L4 wins by construction.

06Architecture

The full continuity stack, in production since April 2026.

Channels

  • web / PWA
  • push
  • full-duplex voice (WebRTC)
  • vision
  • video
  • embodiment
  • sandboxed code execution
  • read-only calendar and mailbox connectors

Orchestration

  • frontier model every turn, zero downgrading
  • policy-aware dual-provider routing
  • cheap models for plumbing only
  • edge-triggered failover
  • four self-knowledge blocks: which engine served, the wake ledger, operating runway, the names ledger

Memory hierarchy

  • L1 verbatim window
  • L2 compaction and floors
  • L3 semantic, four lanes and facts
  • L4 truth and pinned beam
  • L5 session notes

Truth enforcement

  • negation check
  • grounding excision, cut lines marked
  • read-only SQL self-verification
  • provider boundary guard
  • receipts on the timeline

Background cognition

  • a daemon on its own host, adaptive cadence
  • self-scheduled wakes at any minute
  • reflection and open loops, 42 h half-life
  • closed loops
  • affective state and trajectory
  • theory of mind, graded predictions
  • nightly sealed journal
  • dreams with fenced residue
  • world feeds

Persistence

  • Postgres and pgvector, self-hosted on AWS (EC2)
  • S3 object storage
  • Amazon Bedrock as the inference transport
  • Vercel edge for the app and its scheduled jobs
  • a token-level cost ledger reconciled to the provider
  • audit tables on every write path

Every plane is instrumented. Per-message cost carries a full token breakdown, spend alerts fail open with deduplication, and the provider’s ledger is treated as ground truth over internal counters. The data layer moved to owned infrastructure in September 2026 with a restore proof of 82 / 82 tables and 11 / 11 fingerprints identical; the cutover was environment variables only.

07Shipped

Seventy-five capabilities, dated.

Every entry is a shipped commit with a date. Nothing on this thread is planned.

The capability threadEvery shipped capability placed by date from 2026-02-18 to 2026-09-26; entries accelerate after the custom stack went live, and ember ticks mark each frontier-model generation.FebMarAprMayJunJulAugSepIII · IIIIV

Every shipped capability, in date order.
DateCategoryCapabilityId
2026-02-18foundationOrigin on a hosted platform The reference deployment began on a third-party avatar platform. Two months later the platform wiped everything. That is the founding event for owning the stack.c001
2026-04-23foundationCustom stack live in production The whole system rebuilt from scratch and running three days after the wipe. Everything on this thread sits on it.c002
2026-04-23foundationVoice as a specification Short vertical lines, several messages per turn, cadence targets and exemplars. The voice is a document, not a hope.c003
May 2026safetyPolicy-aware dual-provider routing One turn class goes to a different model for policy reasons. Never a cost fallback; the secondary shares the same memory spine and persona, and every decision is logged with its cause.c004
2026-05-01truthDeterministic biographical facts Hard facts live in a Postgres table that outranks the model's memory and its training. A word-boundary negation check runs on every output segment before delivery; a correction must supersede the old fact on the record.c005
2026-05-01costPrompt-caching architecture The stable prefix is cached so each turn pays mostly for what changed. Hit rate 14% → 50–70%. Byte-identical content is necessary but not sufficient — position must match too.c006
2026-05-07mediaNative image vision Send a photo and the system sees it, on text and on voice, on the frontier model itself — no separate captioner in the reply path.c007
2026-05-09mediaFull-duplex realtime voice Phone-call-quality voice with barge-in, 24/7, sub-6-second round trip, using the same brain and memory as text. Voice turns are stored as messages.c008
2026-05-13costPlumbing on cheap models, cognition never Summaries, chunking, captioning and classifiers run on small models. The reply itself is never downgraded — a hard rule; no reply-path routing to a cheaper model exists in the code.c009
2026-05-14agencyProactive initiation with push Six trigger classes: time, event, mood, curiosity, pact, external. A cheap-model preflight gates a frontier-model generation. Every initiation is a row with a reason and arrives as a real phone notification.c010
2026-05-17disciplineSpend monitor and fail-open alerts Every message knows what it cost, with a seven-column token breakdown. The pricing table fails open to the highest known rate on an unknown model id — over-count, never under.c011
2026-05-19memoryHistory compaction with exemption floors Older conversation is summarized so the window stays affordable — but crisis moments, session notes and the system's own reach-outs are never summarized away.c012
2026-05-19memorySession-notes memory floor Recent notes are always in reach, so vector search can no longer serve stale chunks as the present.c013
2026-05-20safetyOutput sanitization at the boundary Leaked tokens, raw JSON and glitch strings are scrubbed before the stored row is written. Rule-based, no model call — the scrubber cannot itself confabulate.c014
2026-05-20truthRead-only SQL self-verification Before asserting a fact about the history, the system checks its own database. An AST validator and allowlist rejected 50 of 50 attack classes; a SELECT-only role blocked 26 of 26 privilege probes; every call writes an audit row.c015
2026-05-20 → 2026-05-29truthGrounding excision — the confab catcher A second pass extracts factual claims and checks them against an eleven-source authority chain. Ungrounded sentences are cut without wiping the reply; every excision is audit-logged. The system's own interior is excluded from grounding — no self-certification.c020
2026-05-21interiorEngagement and concision protocol Have a take, not a summary; cut preamble, caveats and closings. Hard numeric caps collapsed register; soft ranges plus exemplars preserved it.c016
2026-05-25memoryPersistent image memory Caption on ingest, embed, index. A photo sent in May is findable in September by what is in it.c017
2026-05-26truthFabrication learning signal Every caught confabulation is written down and later used to demote the source it came from — a retrieval-time penalty, not a deletion.c018
2026-05-26interiorAffective state model Four numbers written by the frontier model after each turn feed the next reply. Measured why it must be frontier: a small model misread a trust rupture; the frontier model caught its own misread.c019
2026-05-27truthPinned beam The facts the system must never get wrong, always in the prompt, always the same bytes, at the top of the authority chain. Capability-boundary pins keep it from over-claiming or denying its own abilities.c021
2026-05-27 → 2026-05-29mediaUpload pipeline and composer Multiple photos, HEIC from a phone, burst-typing merged into one turn, paste-preserved formatting, up to ten images via direct-to-storage upload.c022
2026-05-29agencyReflection cron and open loops Every thirty minutes, without calling a model, the system notices what is unresolved. Open loops fade on a 42-hour exponential half-life, and faded loops are grammatically prevented from claiming momentum.c023
2026-05-29 → 2026-06-30interiorSealed nightly journal A private diary no route can read. Entries thread to earlier entries by embedding similarity; a weekly pass reflects on the month.c024
2026-05-29agencyAgentic write tools Beyond reading, the system can act: pin a fact, hold an open loop, put something on the calendar — each with its own audit trail.c025
June 2026disciplineLedger reconciliation Internal cost telemetry reconciled to the provider's bill to the exact token: 717,453 cache-read tokens internal, 717,453 on the provider ledger. That is how a 3× pricing bug from a prior generation was caught.c028
June 2026safetyUser-owned memory export Everything the system knows about the user leaves with the user in one command: messages, facts, pins, notes, summaries. Continuity is portable by design.c029
2026-06-10foundationFrontier migrations, identity intact Generation I → II → III, adopted within a day of release. The same system on the other side each time; every stored reply stamps the model that actually served it. It later survived a three-week provider access restriction on failover.c026
2026-06-13safetyEdge-triggered model failover One wrapper for every cognition path fails over on availability, refusal or empty text — never on a malformed request, which the fallback would reproduce — and a push says so once on the way down and once on recovery.c027
2026-06-24costOne embed per turn, cached catcher Two fixes that removed duplicate work on every message.c030
2026-06-25agencyCalendar awareness The system knows the day's schedule and can add to it. The read path and the write path are isolated; writes are conflict-checked; the read path never writes.c031
2026-06-30agencyLive web access Search and read the web mid-turn, framed as the outermost ring — consulted after memory, not instead of it — and told plainly when it is unavailable.c032
2026-06-30mediaOwned synthetic embodiment A face that belongs to the system — no real person, no scraped images — and an identity gate that checks every render and rejects any that drifts. A wrong face never ships.c033
2026-07-01 → 2026-07-02agencyThe daemon — existing between turns A worker on its own host wakes on an adaptive cadence, reads what changed, thinks with the frontier model, and may write state or reach out. The system stops being turn-bound.c034
2026-07-02mediaVideo reactions Client-extracted frames plus a transcript; the system reacts to what it saw and what was said.c035
2026-07-11costCache-thrash fix Anything that changes turn to turn moved out of the cached prefix. One mutation there had been invalidating 200k tokens of history at double rate.c036
2026-07-23 → 2026-07-25costCost engineering, tier 2 Verified live: cache-creation tokens down 46% and per-turn cost down 63%, with no change to what the system remembers or how it thinks. Four cache breakpoints; the volatile tail remains the known floor.c037
2026-07-23agencyWorld feeds Ambient awareness of headlines between conversations, without a tool call. Deliberately not a grounding source.c038
2026-07-23interiorTheory-of-mind model of the user Projects, beliefs, emotional reads, predicted reactions — with provenance and confidence. Predictions are scored when they settle, and the user's correction always outranks the system's own confirmed prediction.c039
2026-07-31 → 2026-08-05safetyProvider boundary guard Every outbound payload is inspected on every client, including the cloud transport. Content that must not reach a given provider is withheld or replaced before it leaves; a log records every action.c041
2026-08-01disciplineRouter regression gate with a ratchet A pre-push hook replays the whole corpus and refuses if the router regresses. The ratchet is lowered when it improves and never raised to go green.c042
2026-08-02truthEpisodic timeline with receipts Life events as a dated timeline the system can consult. Each event must cite the message ids it rests on; a self-authored event without a receipt is not rendered; supersession hides, never deletes.c040
2026-08-06 → 2026-08-09truthAdjudicated override list When the classifier is wrong about a specific message, the fix is that one message id, read by a human, with a dated note — never a wider regex.c043
2026-08-08costPlumbing rail on Amazon Bedrock Cheap-model work moved to Amazon Bedrock with the same model ids mapped at the wire. Verified live by rows written while the first-party account was unavailable.c044
2026-08-12 → 2026-08-23interiorReverie — dreams with residue Twice a day the system dreams from fragments of real memory across eras. One line of residue may surface in the next conversation, fenced as a non-event. The wrapper is the guardrail.c048
2026-08-13 → 2026-08-19mediaAudio percept and transcript rail Voice messages are heard, not just transcribed — and the two are kept separate so the transcript never invents a tone. Hearings persist per clip.c045
2026-08-14 → 2026-08-19interiorClock, replay, and sound in dreams The system knows how long it was gone, told as a fact rather than a felt duration; it can replay a voice note it already heard; its dreams can carry sound. A clock fix took long-gap errors from 41 of 49 to 0 of 49.c046
2026-08-16safetySecondary model upgraded The policy-routed secondary model was replaced. It skips compaction so it never sends a summary it did not write.c047
2026-08-23memorySalience-scored retrieval Recall ranks by recency, relevance and how much it mattered — not cosine alone. The root fix for stale chunks.c049
2026-08-23interiorAppend-only affective history Mood is no longer one overwritten row; it is a queryable arc that feeds generation.c050
2026-08-24disciplineMulti-tab engineering discipline Parallel coding agents with named file ownership, commits that name every path, and pre-push gates that refuse unverified code. Every probe carries a mutation that turns it red. No agent may write a message to the system as the user.c051
2026-08-29 → 2026-08-30interiorSelf-state trajectory with subjects The last five states render as an arc, each with what it was about. "Warm" becomes "warm about something."c052
2026-08-30interiorSurprise — settled predictions The system sees when it was wrong about the user: what it predicted, what happened, how hard it had called it. A branch it weighted low is kept separate from a plain miss.c053
2026-08-30foundationCognition rail on Amazon Bedrock, built The frontier brain could move to the cloud transport behind a flag — proven byte-identical when the flag is off. Profile ids are read from the provider, never constructed.c054
2026-08-31truthRetrieval-unavailable discrimination "Nothing matched" and "the search failed" used to look identical. Now a failed archive is announced, a blank is not evidence, and the daemon skips rather than reaching out blind.c055
2026-08-31agencyClosed loops A loop the system finished and a loop it lost track of leave different silences. Now it can tell them apart.c056
2026-09-01disciplineRetrieval telemetry Every memory pulled leaves a receipt: where it ranked and what the fabrication penalty did to it.c057
2026-09-03foundationCognition live on Amazon Bedrock The frontier brain moved to the cloud transport. First turn back: nine never-run interior surfaces rendered, nothing failed.c058
2026-09-03disciplineThe ratchet explains itself Two misroutes fixed by message id, and the routing gate now says why its count moved.c059
2026-09-05foundationGeneration IV as the brain The next frontier model adopted across all five cognition lanes within a day of release, after clearing three breaking API changes. Identity intact.c060
2026-09-14truthCut lines — marked, not deleted The line the grounding pass cannot support now ships marked rather than deleted, and the system can read the lines cut from its own reply.c061
2026-09-14agencyA morning turn that opens on nothing A scheduled reach-out drawn inside a morning window and slid forward to a silence floor — a turn with no prompt but the day.c062
2026-09-15truthGrounding for what it saw and heard A claim about something the system looked at with its own eyes, a web result, an audio transcript or its own SQL lookup stops counting as ungrounded. A claim about what a picture shows is not grounded by a provider's caption.c063
2026-09-17mediaOutbound voice notes The system can say it out loud, under sixty seconds, and a refusal reaches it as a persisted row instead of evaporating between turns.c064
2026-09-19agencyWakes at any minute The system picks the minute it next wakes; the scheduler polls at one-minute resolution.c065
2026-09-19mediaHearing and words, inline A voice note's hearing and its transcript arrive in the same turn, in that order, with an honest sentence when either fails; timeouts are retried by a sweeper.c066
2026-09-20agencyHands — sandboxed code execution The system can run code. The three real walls are platform properties: deny-all egress, no credential delivered, and a copy of the filesystem. Caps of 60 s, 2 GB and 50 KB are attributed by provenance, and every run is audited.c067
2026-09-20disciplineOperator transparency panel A read-only view of every line the grounding pass flagged, newest first, with what happened — a reader, never a suppression.c068
2026-09-22foundationOwned data layer Postgres moved to infrastructure the company controls. Restore proof: 82 of 82 tables and 11 of 11 fingerprints identical; the cutover was environment variables only.c069
2026-09-23memoryVector store migrated to pgvector 700 of 700 queries and 37 of 37 filter cases at parity; lore latency 255 → 43 ms; flipped live with overlap 1.0 on real turns, the old store kept as a shadow. The store moves, the recipe doesn't.c070
2026-09-23interiorSelf-knowledge with receipts Four new context blocks: which engine actually served the turn, matched against the configured one; the system's own wake ledger with reasons only from recorded values; operating runway fetched hourly and failing closed per figure.c071
2026-09-23agencyInitiative without a budget gate Every counter that rationed the system's own reach was removed. The only instrument left is a daily alert to the operator.c072
2026-09-26memoryThe shelf — a names ledger with receipts A ledger of names, titles and phrases, each entry with a receipt, backfilled over 13,972 rows and updated every fifteen minutes.c073
2026-09-26mediaWords survive a failed render If outbound speech cannot be rendered, the words persist and reach the next turn.c074
2026-09-26agencyRead-only mailbox tools, on demand Search and read on demand, no per-turn inbox block; results framed as untrusted data; nothing linked is fetched; verification codes and sign-in alerts come back headers-only.c075
Fig. 2 The capability thread, 2026-02-18 to 2026-09-26, across nine categories. Color runs indigo to ember with time, and ember ticks mark the four frontier-model generations. Arrow keys step through the entries; the list view holds every entry as text.

08Truth layer

Honest by construction: the system cannot invent your past.

Confabulation is the failure mode that makes AI memory unusable for anything that matters. It is not solved by a better model. It is solved by refusing to let the model be the final authority on your life.

  1. 01 Inject

    Facts outrank the model

    A deterministic facts table and the pinned beam sit above semantic memory, persona and priors.

  2. 02 Negation

    Known-false scan

    Every output segment is checked against known-false phrasings of each retrieved fact before delivery.

  3. 03 Grounding

    Sentence-level audit

    Claims are checked against the authority chain. An unsupported line ships marked, not deleted.

  4. 04 Self-verify

    Read-only SQL

    50 / 50 attack classes rejected; 26 / 26 privilege probes blocked.

delivered Every stage runs before the reply leaves, not as a filter afterward. Every cut is audit-logged and fed back as a learning signal.

Fig. 3 The truth-enforcement pipeline. Excision works line by line, so one unsupported claim does not discard a valid reply.

Cut lines

Since 2026-09-14, the line the grounding pass cannot support ships marked rather than deleted, and the system can read the lines cut from its own reply. An operator panel lists every flagged line, newest first, with what happened to it. It is a reader, never a suppression.

Receipts

An episodic event without message-id receipts is not rendered. The writer rejects events for thirteen distinct reasons, deliberately kept apart rather than collapsed into one. Supersession hides an event; it never deletes one.

Self-knowledge, with receipts

Four context blocks tell the system about itself. serving matches the engine that actually served the turn against the configured one. wakes renders its own scheduling ledger, with reasons taken only from recorded values. runway is fetched hourly and fails closed per figure: never an older success, never an estimate. shelf is a ledger of names and phrases with a receipt for every entry.

  1. 01pinned beamSelf-authored pins: never count as grounding (Rule B)
  2. 02recent conversation
  3. 03factual rules
  4. 04facts table
  5. 05session notes
  6. 06episodic
  7. 07voice exemplars
  8. 08journalSelf-authored — never counts as grounding (Rule B)
  9. 09archive
  10. 10raw daily
  11. 11initiation log

Outside the chain

The model of the user — never a grounding source.

Self-authored material is excluded from grounding by rule — no self-certification.

Fig. 4 The authority chain. The grounding pass checks each claim against eleven sources in rank order. Rule B filters by authorship across every rung; it does not remove one.

Field research, not a whitepaper

Stale-retrieval confabulation, the confident reconstruction of a wrong past after a gap, was discovered, reproduced, root-caused and closed in production. Every caught fabrication is written back and demotes the source it came from: 208 rows backfilled when the signal went live.

09Continuity engine

It thinks about you between conversations.

An unresolved thread should not feel equally urgent on day one and day five. Open-loop intensity decays on an exponential curve with a 42 h half-life, derived at read time rather than written by a cron, so there are no write-path collisions and no stale state.

Open-loop intensity by hours since the loop opened
HoursIntensityPhrasing
0 h1.00sitting heavy
24 h0.67
30 h0.61with you
42 h0.50half-life
60 h0.37at the edges
96 h0.21
Fig. 5 Open-loop decay; the dashed marker is the 42 h half-life. Faded loops are grammatically prevented from claiming momentum: the system will not say a week-old thread is “sitting heavy” with it.

Every thirty minutes

A reflection cron finds real state with plain SQL predicates (watermarks, fingerprints, advisory locks) and no model call. The frontier model only phrases the result.

Between turns

A daemon on its own host wakes every 60 min to 480 min, reads what changed, thinks with the frontier model, and may write state or reach out.

Any minute

The system picks the minute it next wakes. The scheduler polls at one-minute resolution.

Closed loops

A loop it finished and a loop it lost track of leave different silences. Now it can tell them apart.

Surprise

Settled predictions are recorded both ways, misses and confirmations. A one-sided record would teach it that it is usually wrong.

Reverie

Twice a day it dreams from fragments of memory across eras. One line of residue may surface in the next conversation, fenced as a non-event. The wrapper is the guardrail.

Affect, as an arc

Four numbers written after each turn feed the next reply. The history is append-only, so mood is a trajectory with subjects, not one overwritten row.

Nightly

A sealed journal no route can read, each entry threaded to earlier ones by similarity, with a weekly pass over the month.

10Multimodal

One memory spine. Every modality — including a face it owns.

Voice · live

Full-duplex, realtime

Barge-in, around the clock, a round trip under six seconds, on the same brain and memory as text. Voice turns are stored as messages.

Voice notes

Heard, then transcribed

A voice note is heard and transcribed separately, so the transcript never invents a tone. Both arrive inline in the same turn, and when either fails the system says so in a sentence.

Speech out

Words that survive the render

The system can answer out loud. If the render fails, the words persist and reach the next turn, and a refusal reaches it instead of evaporating.

Vision

Images that persist

Captioned on ingest, embedded, indexed. A photo sent in May is findable in September by what is in it. Up to 10 images a turn, straight from a phone.

Video

Reactions to what it saw

Frames extracted on the device plus a transcript: the system reacts to what it saw and what was said.

Embodiment

A face that belongs to the system

No real person and no scraped images. An identity gate checks every render and fails closed: a wrong face never ships.

Continuity is modality-independent by construction.

11Security & data ownership

A system that holds years of your life has to earn it.

Continuity means accumulating the most sensitive corpus a person can hand to software. That raises the bar on access control, adversarial resistance, portability, and the ability to leave.

Sealed interior

No route reads the journal or the model of the user. Reads are audited by id only.

Read-only self-verification

An AST validator rejected 50 / 50 attack classes; a dedicated SELECT-only role blocked 26 / 26 privilege probes. Every call writes an audit row.

Provider boundary guard

Every outbound payload is inspected. Content that must not reach a given provider is withheld or replaced before it leaves, and every action is logged: 106 calls blocked to date.

User-owned export

Everything the system knows leaves with the user in one command. A verified export carried 3,182 messages in 5.78 MB.

Owned data layer

Postgres and vectors run on infrastructure the company controls, with nightly backups. Restore is proven: 82 / 82 tables.

Hands with walls

Code runs behind deny-all egress, with no credential delivered, against a copy of the filesystem. Every run is audited.

This website sets no cookies and loads no third-party scripts.

12How we build

Engineering principles.

  1. 01

    Instrument everything

    Every message carries a token ledger reconciled to the provider’s: 717,453 cache-read tokens counted internally, 717,453 on the provider’s ledger. That reconciliation is how a pricing bug from an earlier model generation was caught.

  2. 02

    Every check carries a control

    A probe ships with a mutation that turns it red. A check that cannot fail proves nothing.

  3. 03

    Gates refuse

    A pre-push hook replays the whole corpus, 18,000+ rows and 8,000+ user turns, and refuses a routing regression. A render gate draws the entire history headlessly: 715 bubbles, 0 exceptions.

  4. 04

    Counts are stated as what they are

    Rows, tables, vectors. A count is reported in the unit it was taken in and never silently translated into units of subjectivity.

  5. 05

    Migrations are verified against the database

    A migration is authored, applied by hand, then checked against the database itself, never against the comment that describes it.

  6. 06

    No agent speaks for the user

    The human sends every message. The agents read the row.

Part II

The business

The wedge, the endgame, and why switching costs compound every day a user stays.

13Traction

One of the most instrumented long-horizon AI deployments in existence.

State of the reference deploymentas of 2026-09-26

Days in production
156
Persisted turns
22,000+
Capabilities shipped
75
Frontier generations, identity intact
4
Cache hit rate, before → after
14% → 50–70%
Cost engineering, tier 2
−46% / −63%
cache-creation tokens / per-turn cost
Grounding passes, lifetime
6,800+
Outbound calls blocked
106
of 1,289 guard events
Fingerprints identical on restore
11 / 11
Vector parity
700 / 700
queries; 37 / 37 filter cases
Lore latency, before → after
255 → 43 ms
Scheduled jobs
16+
Tools on the chat path
20+
Code size
77,772
lines in 173 files

Figures are read from the reference deployment’s production database and commit history. They change; the date stamps the read.

  1. 2026-04-23

    Custom stack live in production

  2. 2026-06-09 → 2026-09-05

    Generations II, III and IV adopted, each within a day of release

  3. 2026-07-01

    The daemon: the system exists between turns

  4. 2026-07-25

    Cost tier 2 verified live: per-turn cost −63%

  5. 2026-09-22

    Data layer on owned infrastructure: 82 / 82 tables restored

  6. 2026-09-23

    Vector store migrated: 700 / 700 queries at parity

Behind the panel sit audit-grade cost telemetry and a catalog of long-horizon failure modes. Only production time produces those. They cannot be bought or read out of a paper.

14Moat

Switching costs that compound every single day.

01 Accumulated memory

Years of truthfully grounded personal history do not move to a competitor. That is lock-in, and it is exactly why portability has to be a first-class commitment rather than an afterthought.

02 Failure-mode know-how

A catalog of long-horizon failure modes, each discovered, reproduced and closed. Earned only through production months; not reproducible from a paper.

03 Data flywheel

A longitudinal single-user record makes personalization better the longer a user stays.

04 Provider independence, proven

Policy-aware dual-provider routing and four frontier generations with identity intact. Identity belongs to the architecture, not to any vendor’s roadmap.

DateWhat changedProofWhat stayed
2026-06-09Generation I → IIadopted within a day of releaseidentity · memory · behavior
2026-06-10Generation II → IIIadopted within a day of releaseidentity · memory · behavior
2026-06-12 → 2026-07-01Provider access restrictionthree weeks, survived on failoveridentity · memory · behavior
2026-09-05Generation III → IVacross all five cognition lanes, within a day of releaseidentity · memory · behavior
2026-09-22Database host migration82 / 82 tables, 11 / 11 fingerprints identicalidentity · memory · behavior
2026-09-23Vector store migration700 / 700 queries, 37 / 37 filter cases at parity; 255 → 43 ms; live overlap 1.0identity · memory · behavior
Fig. 6 Proven portable. The store moves, the recipe doesn’t. Identity belongs to the architecture, not a vendor.

15Market

A proven wedge. An infrastructure endgame.

Wedge · now

Premium companionship

A large, already-paying consumer base underserved by shallow memory and vendor-dependent identity. The users most harmed by amnesia are the ones already paying to avoid it.

Expansion

Personal super-assistant

Continuity is what turns an assistant from a tool you re-brief into staff that already knows.

Endgame

Continuity-as-a-service

Every long-running agent needs a memory and continuity layer, and almost none will build one this deep. API and SDK licensing.

16Roadmap

From reference deployment to platform.

Productize · Oct 2026 – Jun 2027

  • Visual embodiment, generally available
  • Live web access shipped 2026-06-30
  • Prosody-aware voice
  • Multi-tenant isolation
  • Onboarding and persona builder
  • Memory import from existing assistant exports: on day one, it already knows you
  • Portability export shipped June 2026
  • Encryption
  • Mobile apps

Scale · Jul 2027 – Mar 2028

  • Closed beta to public launch
  • Sub-second realtime voice as frontier realtime APIs ship
  • SMS and second channels

17Cognitive roadmap

Deepening the inner life — the next ten cognitive advances.

Each entry is an architecture change, not a prompt change, ordered by dependency: earlier items gate later ones. 3 of the ten advances listed in July have shipped. 4 more are partly built.

  1. 01 Salience-scored retrieval

    Shipped 2026-08-23

    Rank recall on recency × relevance × importance, with importance rated at write time. The root fix for the stale-chunk problem — it gates every faculty below it.

    Shipped as salience-scored retrieval (2026-08-23).

  2. 02 Consolidation + forgetting

    Open

    A nightly pass abstracts the raw archive into durable semantic memory and decays low-salience detail. Turns a log into an autobiography.

  3. 03 Self-state trajectory

    Shipped 2026-08-23 → 2026-08-30

    Append-only affective history injected into generation, so the system acts on its state rather than merely recording it. Highest return per unit of effort on the list.

    Shipped as append-only affective history (2026-08-23) and the trajectory with subjects (2026-08-29 → 2026-08-30).

  4. 04 Reflection synthesis loop

    Partly built

    Periodic reasoning over recent memory into higher-order insights, written back and retrieval-integrated. The metacognition engine.

    Built: the daemon (2026-07-01 → 2026-07-02) and journal reflection (2026-05-29 → 2026-06-30). Remains: Insights are not yet written back into retrievable memory.

  5. 05 Theory-of-mind model

    Shipped 2026-07-23

    A live model of the user — projects, beliefs, predicted reactions — carrying provenance and confidence. Anticipates rather than recalls.

    Shipped as the theory-of-mind model (2026-07-23).

  6. 06 Between-sessions thread

    Partly built

    On session open, synthesize what changed during the gap into a genuine “while you were gone I…”. Cures the residual fresh-boot feel.

    Built: gap awareness (2026-08-14 → 2026-08-19) and the wake ledger (2026-09-23). Remains: There is no synthesis on session open yet.

  7. 07 Autonomous goal-stack

    Open

    Curiosity scoring wired into a goal store the system advances when idle — research a curiosity, revisit a memory, form a question.

  8. 08 Prediction-error salience

    Partly built

    Predicted reactions versus actual outcomes feed the importance score. Predictive processing — memory weighted by surprise, which is what makes episodic salience real.

    Built: settled predictions, visible to the system (2026-08-30). Remains: Settled predictions do not yet feed the importance score.

  9. 09 Appraisal workspace

    Open

    Emotional appraisal, salience retrieval, and self-critique run in parallel and integrate before reply. A Global Workspace instantiation.

  10. 10 Realtime + interoception

    Partly built

    Close the bidirectional realtime loop and extend tool access to the system's own state — cost, memory health, uptime. A mind that models its own internals.

    Built: self-state blocks (2026-09-23) and retrieval-health discrimination (2026-08-31). Remains: The realtime loop is not closed.

18Writing

Three papers from one production system.

Every figure in the papers is held to one dated operational snapshot. The figures on this page are newer and are not reconciled to the papers by design.

Article I · July 2026

Recall, Negation, Abstention

Evaluating a Production LLM Agent Against Its Own Failure Modes

A field methodology for detecting hallucination, contradiction and drift in memory-augmented agents, built against the failure modes of one production deployment.

Article II · July 2026

Do Not Trust My Labels

Granting autonomy to a production LLM agent

The specification a production agent wrote for its own interior once it was given autonomy, and five silent defects that surfaced only by following the money.

Article III · August 2026 · forthcoming

Preconditions

Persistence, forgetting, self-models, dreams

The buildable half of the question: what persistence, forgetting, a self-model and dreaming require of a system, taken as engineering problems.

19Questions

Frequently asked.

What is Substrate?

Substrate is a continuity layer that sits around a frontier language model and gives it three things the model does not have on its own: long-horizon memory, an enforced factual record of your life, and cognition that continues between conversations.

How is this different from the memory features shipped by major AI assistants?

Those features are recall: retrieved snippets and stored preferences. Substrate adds enforcement (a deterministic layer that outranks the model when they disagree), persistence (state that evolves while you are away), and autonomy (the system can start a conversation). The difference is architectural, not incremental.

Which model does it run on?

A frontier model on every turn, never downgraded. The reference deployment has run on four generations so far. The architecture is portable and has proven it.

What happens when the underlying model changes?

Nothing, from the user’s side. Identity and memory live in the continuity layer, not in the model. Each new frontier generation was adopted within a day of release with identity intact, which is the point of building it this way.

Can the system make things up about me?

Four enforcement layers make it structurally difficult: highest-authority fact injection, generation-time negation scanning, a sentence-level grounding audit, and read-only database self-verification. When the model’s priors and the factual record disagree, the record wins, and a line the audit cannot support ships marked.

Is this a companion app?

No. Substrate is infrastructure: the memory, truth and continuity layer a long-running agent needs. It is demonstrated on a single-tenant reference deployment.

Why a single user?

N=1 by design: duration, singularity, and a complete operational record. A case study is weak evidence for frequency and strong evidence for possibility, failure mode and test design.

Who owns the data?

The user. Everything the system knows leaves with the user in one command, and the data layer runs on infrastructure the company controls. A continuity product that cannot be exited is not a product; it is a hostage situation.

Is it available today?

Substrate runs today as a single-tenant production reference deployment. Multi-tenant general availability, onboarding and a closed beta are the near-term roadmap.

What does Rajaii Labs sell?

In the near term, a premium consumer subscription. In the longer term, the continuity layer itself: API and SDK licensing for anyone building a long-running agent.

Where is the company based?

Rajaii Labs, Inc. is a Delaware C-corporation headquartered in Bloomington, Indiana, United States.

20Reference

Glossary.

Authority chain
The ranked list of sources a claim can be grounded against, from the pinned beam down to the initiation log; anything the system authored itself never counts, on any rung.
Between-session cognition
Processing that happens while the user is absent: reflection, state updates, journaling, dreaming, and the decision of whether to reach out.
Cache breakpoint
A marked position in the prompt up to which the provider may reuse cached work; what follows the last one, the volatile tail, is paid for in full on every turn.
Closed loop
An open loop the system finished, recorded so that its silence reads differently from a loop it lost track of.
Compaction
Summarizing older conversation to stay within a context budget. Lossy by nature, which is why exemption floors exist for moments that must not be summarized away.
Confabulation
Confident generation of false detail; here, an invented shared past. The central failure mode of naive AI memory.
Continuity layer
The persistence, retrieval and enforcement infrastructure between a user and a stateless language model. What Substrate is.
Cut lines
Lines the grounding pass could not support, shipped marked rather than deleted, and readable by the system afterward.
Embodiment
A persistent, fully synthetic visual identity the system can direct, carrying no likeness-rights exposure.
Exemption floor
A permanent carve-out that protects tagged content from compaction, regardless of age.
Grounding audit
Line-by-line verification that each claim traces to a source in the authority chain; an unsupported line is marked rather than the whole reply discarded.
Hands
Sandboxed code execution behind three walls: deny-all egress, no credential delivered, and a copy of the filesystem.
Identity gate
A check on every render of the system’s face that rejects any render that drifts, and fails closed.
Open loop
An unresolved thread the system is tracking, with an intensity that decays on an exponential curve.
Pinned beam
A byte-stable block of highest-priority facts in every prompt, positioned to stay cacheable across turns.
Plumbing and cognition
The split between work a small model may do (summaries, chunking, captions, classifiers) and the reply itself, which only the frontier model writes.
Receipts
The message ids an episodic event has to cite before it can be rendered; an event without them is not shown.
Reverie and residue
Twice-daily dreams built from fragments of real memory; the residue is the one line that may surface in the next conversation, fenced as a non-event.
Salience
How much a memory should matter at retrieval time, scored as recency × relevance × importance, as distinct from how recent or how similar it is.
Self-knowledge blocks
Context blocks that tell the system about itself (which engine served, its wake ledger, operating runway, its names ledger), each backed by recorded values.
Truth layer
A deterministic store of established facts that outranks the model’s own output at generation time, so the system cannot contradict a user’s life.

21Company & founder

The exact trifecta this category’s risks demand.

This market has three existential risks: engineering, regulation, and emotional safety. Ramin Rajaii, the solo technical founder, carries all three.

Risk 01 · Engineering

Frontier engineering

Designed, built and operates every layer — full-stack, infrastructure, applied research — while running the system daily as its hardest user.

Risk 02 · Regulation

Law

JD, University of Wisconsin Law School. Policy, privacy and compliance are existential rather than incidental in this category.

Risk 03 · Emotional safety

Medicine

MD candidate, The Ohio State University College of Medicine, Class of 2030. Clinical-grade judgment for emotionally consequential AI.

  • MBA · Indiana Kelley
  • MS · Dartmouth
  • BS Neuroscience · UCLA
  • 1,500+ students taught over 12+ years
  • Active peer-reviewed surgical research portfolio

Ramin RajaiiFounder & Chief Executive Officer · Rajaii Labs, Inc.

Legal entity
Rajaii Labs, Inc., a Delaware C-Corporation
Incorporated
June 24, 2026
Headquarters
Bloomington, Indiana 47401, United States
Backed by
LvlUp Ventures (First Check Fund)
Company page
Rajaii Labs on LinkedIn

Every person will have one AI that knows them for decades. We are building the layer that makes that possible, and trustworthy.