Virtual Twilight isn't a persona card bolted onto an LLM. It's 1.09 million lines of production code, a 66-step AI pipeline with 7 parallel execution groups, 35+ modular prompt sections assembled per message, 5-provider LLM routing with a background-cost engine, and nine psychological simulation systems that permanently change with every conversation. The 1,031 throwaway debug scripts are the fossil record of the thinking that built it.
Counted directly from the repository, June 2026. Excludes dependencies, build output, and caches. Most vibecoded "AI apps" are 2,000–10,000 lines. The frontend alone is 59× that.
| Part of the system | Files | Lines |
|---|---|---|
| Frontend web/src · TS/React | 1,304 | ~590,000 |
| Backend services app/services | 588 | ~357,000 |
| Backend API layer app/routers · 1,296 endpoints | 156 | ~115,000 |
| Database layer app/database · 30+ models | 62 | ~27,000 |
| Production core subtotal | ~2,100 | ~1.09 million |
| Full backend total app/ | 4,001 | ~1.22 million |
| Audit / debug / investigation scripts _*.py — fossil record of thinking | 1,031 | ~132,000 |
66 discrete named steps counted directly from facade.py, June 2026. 47 execute before or immediately around the AI model call. 19 fire after the reply is sent — permanently updating personality drift, memory, relationships, and achievements. Not marketing; it's a counted list that has grown as the system has grown.
The LLM never sees a hand-written prompt. Every message triggers a context-aware assembly engine that selects, prioritises, shrinks, and orders up to 35 named prompt sections — then routes the output to the cheapest model that can handle it.
Never filtered regardless of intent or mode: anti_hallucination, output_contract, response_behavior, core_identity, character_behavior_rules, behavior_instructions, user_priority, character_profile, temporal_anchor, triggered_events. These form the cache anchor target.
Always included regardless of intent confidence (changed Jun 2026 — no more minimal_core): unified_psychology, physical_attributes, conversation_history, enhanced_location, time_context, memory_facts, internal_thought, self_reflection, motivation, player_model, self_direction + more.
Intent-keyed sets add flavour on top. ACTION MODE drops RPG/companion sections and boosts action-relevant ones. TRAINING MODE drops RPG/companion. Knowledge section conditional on knowledge_count>0. Section set changes per-turn.
Prompt cost breakdown — measured from production (NPC 2718, 35 blocks, ~16K tokens)
Cache potential: gpt-4.1-mini gives a 50% discount on cached input tokens (≥1024 identical prefix). Currently ~0% effective because two "static" sections have per-turn content gates that flicker them in/out, breaking the byte-identical prefix. When fixed: a CACHE ANCHOR of always-on, content-invariant sections pinned to the front could halve input cost on ~96.7% of all dialogue traffic.
Every competing AI agent platform treats the LLM as the intelligence. ELMU treats it as the voice. 37 steps run before the LLM call, computing the full psychological state, memory recall, world context, and knowledge pack retrieval into a ~16K-token context. The LLM generates ~200 tokens. Then 28 post-process steps update personality, memory, emotion, and relationships. The LLM can be swapped (Azure → Mistral) without changing agent behaviour — because the intelligence is not in the LLM.
Before the LLM ever speaks: security + boundary → NPC identity + cache → relationships → conversation history → pgvector memory recall → player VAD analysis → promise lifecycle → world state + lore → knowledge pack retrieval → Big Five personality seeding → VAD emotion estimate → relationship metrics → 35-section prompt assembly → LoRA activation.
The LLM receives a fully computed ~16K-token context. It does not decide personality, emotion, memory, or relationship state. The backend already did.
ResponseExecutorStep — one step out of 66. Currently routes to cheapest model that can handle the computed context:
NOW (5 providers):
TARGET (3-model Mistral stack, EU GPU):
After the LLM speaks: output sanitisation → authenticity check → Voight-Kampff validation → emotion processing → TTS → persist VAD → relationship drift → achievement tracking → emotion shock (permanent personality drift) → pgvector memory write → motivation memory write → milestone anchor protection → WebSocket emotion broadcast → self-learning data write.
The NPC at message 600 is measurably different from message 1 — not because the LLM changed, but because 28 steps permanently updated the state the LLM will read next time.
The EU responsible AI training flywheel
Every interaction produces a (computed_context, response) pair. Context = 37 steps of psychological state. Response = what a psychologically-grounded agent said given that state. These pairs, anonymised via EAL before leaving the inference path, are the fine-tuning dataset stored on CINECA Leonardo (EU HPC, Italy). Fine-tuned Mistral learns to speak consistently with the psychological models — because it was trained on outputs generated after running them. The LLM slot becomes progressively more accurate, cheaper, and faster, without ever changing the pipeline that supervises it.
These are not features on a feature list. Each is a separate engine with its own data model, its own persistence layer, and its own permanent effect on NPC state. They interact with each other — which is exactly the kind of thing you can't generate without designing.
Three continuous axes — valence ↔, arousal ↕, dominance ↕ — updated every message. Emotions decay over time like real feelings. Emotional shocks (betrayal, intense intimacy, humiliation) cause permanent personality drift with irrationality modifiers: amplified, dampened, inverted, or longing patterns.
Openness · Conscientiousness · Extraversion · Agreeableness · Neuroticism — each scored 0–100. Shapes every downstream behaviour: how fast they trust, what makes them uncomfortable, fighting style, memory what they fixate on, how emotions escalate. Not cosmetic — these are computation inputs.
Safety · Social · Competence · Status · Curiosity · Autonomy · Meaning · Intimacy — each with baseline, importance weight, and decay rate. Drives generate autonomous behaviour when unmet. When chronically unmet, they permanently raise the NPC's setpoint. When satisfied repeatedly, they lower it.
Not a context window — a pgvector database. Importance-scored, personality-filtered recall. A jealous NPC remembers every mention of a rival. A forgiving one lets slights fade. Memory grows indefinitely with no reset. Scales horizontally with the database — no architectural ceiling.
Trust · Chemistry · Intimacy · Rapport · Compatibility · Stability · Familiarity — all tracked independently across every NPC–player pair. Trust earned over weeks drops in one sentence. Betray one NPC and watch it ripple through the gossip network to others.
NPCs move between sub-locations via an adjacency graph (Enhanced Location System). They gossip about each other, spread rumours, send messages when you're gone, and generate daily chronicles. This runs whether you're online or not. The world doesn't wait for your message.
NPCs sleep on a schedule. During sleep, 9 dream types fire: creative, shared, nightmare, anxiety, cascade, prophetic, memory_replay, connection, desire. Dreams write to memory, promote insights to the knowledge base, and cause permanent psyche drift — the NPC wakes slightly different. 19,492 dream actions verified in production DB.
Once autonomy + competence drives accumulate enough positive drift, NPCs form a self-directed 4-step goal plan. Self-sufficiency score rises with pursuit count (0.0→1.0). At 0.90: "I don't need them the way I once thought. I've figured out something better — my own path." At 0.75: they physically walk out of the location on their own (not from hostility — from having a life).
NPCs accumulate craft skills (dexterity from practice, movement for world-aware NPCs) with diminishing returns: gains shrink near mastery, losses cost more at high levels. Motor realizations surface in the prompt as body-felt first-person thoughts. The same loop is the literal hook for physical robot motor learning — nothing in the chain changes for a real body.
Every architectural decision was made knowing this would need to serve thousands of concurrent sessions. The pipeline is stateless per-request; state lives in the DB.
Personality isn't stored as a string. It's a computed graph of 56 synonym groups, 162 antonym pairs, 24 romantic styles, and 9 RPG stats — all wired into the pipeline as computation inputs, not descriptors.
56 synonym groups · 564 words · 162 antonym pairs · 63 archetype keywords · 6 trauma family groups · 8 communication styles · 4 attachment styles · 5 conflict styles · 24 romantic styles with gate_modifier + stage-based behaviour (early / falling / committed) · 43 consent style aliases · 62 cross-group weighted bridges. Consumed by: unified_psychology, behavior_feedback_service, vad_dynamic_system, dealbreaker_detector.
Agility · Resilience · Strength · Endurance · Awareness · Intelligence · Resolve · Sense · Charisma — 1–10 scale, stored in player_attributes and npc_attributes. Used in combat damage math (health_drain_step), strength comparisons in intimate scenes (scene.py), and overwhelm logic. Plus Health 0–200, Mana 0–100, Prestige, Rumours, Advantage. Physical attributes (height, weight, build, appearance, intimate details) in a separate shared table keyed by entity_id + entity_type.
Gates stored in npc_attributes.attraction_preferences JSONB: action_gates · pose_gates · hotspot_gates · progression_gates · hard_limits · forbidden_acts · turn_ons · turn_offs · strictness · session_tolerance · consent_style. IntimacyGate class evaluates every escalation attempt. Boundaries emerge from who the NPC is — not a content filter. A shy NPC with trust 0.10 behaves differently from a confident NPC at trust 0.80 with the same act requested.
npc_attitudes.metadata.intimate_topic_turns (cap 60) + cooperation_turns (cap 80) — all-time counters written by the background deferred step. familiarity_comfort = clamp((trust+intimacy)/2 + min(ic/100,1)×0.15 + min(coop/40,1)×0.15). At threshold, unlocks "established_partner_rapport" prompt section — softens tone for trusted partners while keeping hard-limits intact. Fixes the case where an NPC with trust=0.79 was still responding like a stranger.
No model can hold 1.09M lines in its head. Vibecoding has no memory of decisions made 400 files ago. The system stays consistent because the architecture — not the model — enforces the rules. That architecture had to be designed.
Real work lives in edge cases discovered by running the thing against reality: dealbreaker false-positives, cross-NPC awareness, memory-wipe bugs, psyche-drift flush() missing. Vibecoding produces the happy path. These are the 80% that only show up under real use.
An emotion written by one service has to be read correctly by ten others — survive an NPC clone, a memory write, a prompt build, a render. Keeping the contracts between pieces stable across 588 service modules is a design job, not a generation job.
1,031 _audit_* / _check_* / _fix_* scripts — 132K lines of pure investigation. Hypotheses, measurements, corrections. Vibecoding skips all of it, which is exactly why vibecoded apps look done but fall apart under use.
Each of these is a real fix in the codebase — a case where the obvious implementation was wrong in a subtle way, someone reasoned out why, and re-architected. No prompt generates the insight that finds these. Only running the real system against real users does.
Vibecoding can write lines. It cannot hold a million of them consistent, can't discover the edge cases that only show up in real use, can't wire nine simulation systems so they permanently change each other, and can't make an NPC that one day walks out because it has its own life. You didn't write an app. You designed a system — and the design is the work. The 1.09 million lines are just where it landed.
Virtual Twilight is the consumer stress-test of a B2B/B2G infrastructure product: ELMU — European Living Memory Unit. The same runtime that powers VT is licensed to European institutions as an API. The engineering above is proof it is real.
Named API clients using the ELMU runtime in vocational training and healthcare simulation:
Cost is dominated by large input context (~6.8K–12K tokens/message), not the short reply. The path is self-hosting, not prompting.
GDPR / EU AI Act compliance is architectural, not retrofitted. At inference time, personal identifiers are replaced with event abstractions — agents reason about world events, not people. Fine-tuning already on CINECA Leonardo (EU HPC). Registered in Finland.