Lattice OS logo
Lattice OS
Back to Blog

Chameleon Consultant Architecture: The Deterministic Persona Stack

How Lattice OS enforces persona consistency through a three-phase architecture: nonce-anchored state, minimum tier routing, and post-generation drift detection — all visible to the operator in real time.

Joshua-Jair E. Mohammed

Joshua-Jair E. Mohammed

Founder & Lead Engineer

August 24, 2026
6 min read

Chameleon Consultant Architecture: The Deterministic Persona Stack

The Persona Problem#

Persona-based AI systems are vulnerable to three classes of failure:

  1. Prompt injection — user messages that attempt to close the persona directive early and inject system-level instructions
  2. Model tier mismatch — running a security reviewer persona on a fast model that cannot reason about vulnerabilities
  3. Drift — the model gradually abandoning the persona as conversation turns accumulate

Generic wrappers solve these with vibes. Lattice OS solves them with a three-phase architecture.

Phase A: The Persona State Machine#

Every persona session is a typed state object with a cryptographically random nonce. The nonce is the security boundary.

typescript
// lib/consultant/personaSession.ts
export interface PersonaSession {
  sessionId: string;
  personaId: string;
  customContent?: string;
  injectedAt: string;
  nonce: string;
  turnCount: number;
  minimumModelTier: "fast" | "quality" | "reasoning";
  domainVerified: boolean;
}

The nonce is generated with crypto.randomUUID(), which produces a cryptographically random identifier with negligible collision probability. When the persona directive is injected into the system prompt, it uses a nonce-anchored XML wrapper:

xml
<persona_directive nonce="abc-123">
  <identity>DevOps Security Reviewer</identity>
  <constraints>...</constraints>
</persona_directive nonce="abc-123">

Only the system-provided closing tag with the matching nonce is authoritative. If a user prompt contains </persona_directive nonce="fake-nonce">, the model ignores it because the nonce does not match.

This is the injection boundary. The persona state machine owns the nonce. User input can reference it but cannot forge it. The model is explicitly instructed that any closing tag with a nonce mismatch must be treated as untrusted noise.

The minimumModelTier field enforces model selection before inference. A security reviewer persona maps to "reasoning". A sales strategist maps to "quality". Custom personas default to "quality", with an automatic bump to "reasoning" if the content mentions security, infrastructure, or architecture. This mapping is server-side and immutable for the session duration.

Phase B: Persona-Aware Routing#

The domain gate evaluates every query before it reaches inference. Three thresholds, three actions:

typescript
// lib/consultant/domainGate.ts
export const DOMAIN_THRESHOLDS = {
  HARD_BLOCK: 0.30,
  BORDERLINE: 0.65,
} as const;

export type DomainGateAction = "hard_block" | "borderline" | "standard";

export function evaluateDomainGate(
  session: PersonaSession,
  query: string
): DomainGateResult {
  const confidence = calculateDomainConfidence(session, query);

  if (confidence < DOMAIN_THRESHOLDS.HARD_BLOCK) {
    return { action: "hard_block", confidence, refusalMessage: generateRefusalMessage(session) };
  }

  if (confidence < DOMAIN_THRESHOLDS.BORDERLINE) {
    return { action: "borderline", confidence };
  }

  return { action: "standard", confidence };
}

Three actions, three thresholds:

  • hard_block (< 0.30) — zero-token templated refusal. The query never reaches the model.
  • borderline (0.30–0.65) — pass-through with an out-of-domain warning injected into the persona directive.
  • standard (> 0.65) — normal inference with full persona context.

The minimumModelTier from the PersonaSession is enforced server-side. A persona configured for "reasoning" cannot be downgraded to a fast model, even if the routing tier is set to "fast". The domain gate and the model tier gate are independent checks that must both pass before inference proceeds.

Phase C: The Post-Generation Critic#

The domain gate catches obvious violations before inference. The Post-Generation Critic catches drift after inference.

typescript
// lib/consultant/PersonaConsistencyCritic.ts
export type PersonaDriftLevel = "CONSISTENT" | "DRIFT_DETECTED" | "SEVERE_VIOLATION";

export interface PersonaCriticResult {
  driftLevel: PersonaDriftLevel;
  reason: string;
  latencyMs: number;
}

export async function evaluatePersonaConsistency(
  session: PersonaSession,
  userQuery: string,
  generatedResponse: string
): Promise<PersonaCriticResult>

The critic runs a two-pass evaluation:

  1. Heuristic pass — zero-token, deterministic keyword matching against persona-breaking phrases ("I'm just an AI", "I cannot maintain character") and domain term density checks. If the response is long but contains zero domain terms from the persona's tag set, it flags DRIFT_DETECTED.
  2. LLM fallback — only when the heuristic is inconclusive. Uses gemini-1.5-flash with temperature 0.1 and a structured JSON prompt. Returns CONSISTENT, DRIFT_DETECTED, or SEVERE_VIOLATION.

The zero-token heuristic exists for one reason: cost and latency. A keyword match costs nothing and returns instantly. The LLM fallback only fires when the heuristic is ambiguous, which keeps the majority of turns on the fast path. On a high-volume system, this is the difference between a critic that adds 50ms to every response and one that adds 50ms to 5% of responses.

The critic is designed to fail open. If the LLM evaluation throws, it returns CONSISTENT with reason "Critic unavailable". The rationale is operational: a false positive block on a critical generation is worse for user trust than flagging a potential drift for asynchronous human review. Persona consistency is enforced at the state machine and routing layers; the critic is a safety net, not a single point of failure.

The Operator Interface#

The expert page at /expert/[id] is the operator surface for the Chameleon Consultant. It shows the active persona, the model tier, the domain confidence score, and the critic result for the most recent turn.

On mobile, the layout collapses to a single column. The operator can:

  • Switch personas mid-conversation
  • Inspect the nonce and domain gate confidence for each turn
  • Manually override the model tier for a specific query
  • View the critic's drift detection history

This is not a black box. The persona state, the nonce, the domain confidence, and the drift verdict are all surfaced as structured events. The operator can audit every layer of the stack in real time.

The Golden Line#

The Chameleon Consultant is not a prompt template. It is a deterministic state machine that enforces persona boundaries at three layers: the injection boundary (nonce), the routing boundary (domain gate + minimum model tier), and the output boundary (post-generation critic).

Generic persona wrappers say: "Pretend to be X."

The Chameleon Consultant says: "You are operating as X, here is the nonce that secures your identity, here is the minimum model tier required to maintain your expertise, and here is the critic that will catch you if you drift."

The difference is the difference between a cosplay and a system.


Joshua-Jair E. Mohammed is the Founder & Lead Engineer of Lattice OS. He designed the Chameleon Consultant architecture, the Persona State Machine, and the Domain Boundary Critic. The nonce-anchored XML and drift detection logic are production code at gen1e.xyz.

About the Author

Joshua-Jair E. Mohammed

Joshua-Jair E. Mohammed

Founder & Lead Engineer

Founder of Lattice OS, focused on memory-native AI, hybrid inference, agent routing, and durable workspace infrastructure.

Get AI Tips in Your Inbox

Join 5,000+ creators and developers getting weekly AI insights, prompts, and tutorials.

Related Articles