The Sirius Document · v2

The Architecture Behind NaviStrat

A complete, plain-English tour of how the platform thinks, remembers, and acts — from the moment a user types a message to the moment a treasury workflow fires off. Diagrams, stick figures, and examples included. No engineering degree required.

01 · Big Picture

The 30-second map

Every NaviStrat experience is a conversation between five layers. A person (on any device) talks to the App Shell, which hands the message to the AI Engine. The engine picks a brain (LLM Provider), reads and writes memories (Data Layer), and reaches out to the outside world through Integrations and Workflows.

🧑 In plain words
The analogy: Think of it as a person (the AI engine) sitting at a desk. The app shell is the receptionist who walks your question over. The LLM provider is the brain the person uses to think. The data layer is their notebook. Integrations are their phone and email. Workflows are their calendar alarms that fire even when they're asleep.
⚙️ Under the hood
The stack: React + Tailwind on Vite for the shell; a single backend function (wingmanChat) orchestrates provider calls; Base44 entities with row-level security hold all state; OAuth connectors (Gmail, Calendar, Drive, Sheets) feed event-driven workflows. Everything runs on the Base44 platform — auth, hosting, DB, and scheduling are managed services.
Youphone + laptop
type
🖥️
App Shell
React + Tailwind
calls
🧠
AI Engine
wingmanChat
routes to
⚡
LLM Providers
OpenAI · Anthropic · Google
reads/writes
🗄️
Data Layer
Entities + RLS
connects to
🔌
Integrations
Gmail · Calendar · Drive · Sheets
triggers
⏰
Workflows
scheduled + event-driven
02 · Core

The AI Engine: how one message becomes one reply

When you press send, the message doesn't go straight to an AI. It first passes through a context tier selector that decides how much "brain" the question deserves. A quick "thanks!" needs almost no context; a request to "draft a board paper from last quarter's cash data" needs the full memory vault and your saved documents.

The engine then assembles a single package: the system prompt (Sirius's personality and rules), the most relevant memories, and any vault documents that might matter. That package is handed to wingmanChat, which calls the chosen provider directly and streams the reply back.

🧑 In plain words
The analogy: You're not shouting into a void. There's a receptionist (tier selector) who decides whether your question needs a quick answer or a committee meeting, a researcher (memory fetch) who pulls the right files, and a speaker (the provider) who actually answers. Each step is cheap unless the question is hard.
⚙️ Under the hood
The flow: useWingmanSync → classifyTier() (from wingmanContextTiers.js) → memory/vault fetch → base44.functions.invoke('wingmanChat', { messages, tier, context_mode }) → provider API → reply. The whole round-trip is one SDK call from the frontend; the backend function owns provider selection, caching, and failover.
Usertypes a message
1. send
🎚️
Context Tier Selector
cheap · mid · premium
2. pick brain size
📋
System Prompt
🧩
Relevant Memories
📚
Vault Docs
3. assemble context
⚙️
wingmanChat (backend)
direct-to-provider · caching · router
4. call provider
⚡
LLM (OpenAI / Anthropic / Google)
5. stream reply back
Userreads answer
03 · Efficiency

Context tiers: paying for the right amount of thinking

Not every message deserves the same brain. A two-word reply shouldn't cost the same as a 2,000-word analysis. The tier selector reads each message and routes it to one of three tiers — each with a different model, a different amount of conversation history, and a different price.

🧑 In plain words
The analogy: It's like a hospital triage. A paper cut gets a nurse (cheap, fast). A broken arm gets a doctor (mid). Heart surgery gets the chief surgeon (premium). You'd never call the surgeon for a paper cut — and the tier system makes sure we don't.
⚙️ Under the hood
The selector: wingmanContextTiers.js classifies by keywords and length: greetings and confirmations → minimal; normal turns → standard; long messages, file attachments, or "analyze/draft/write" verbs → full. The tier becomes the tier field on the wingmanChat call, which maps to a provider + model.
MinimalGemini Flash / Haiku
📦 Context: Last 2 messages💬 When: "thanks", "ok", quick facts💰 ~$0.10 / 1M tokens🎚️ Tier 1
StandardSonnet / GPT-4o
📦 Context: Last 8 msgs + 3 memories💬 When: Most conversations💰 ~$3 / 1M tokens🎚️ Tier 2
FullOpus / GPT-4o
📦 Context: Everything + vault + files💬 When: Long analysis, documents💰 ~$15 / 1M tokens🎚️ Tier 3
04 · The big lever

Prompt caching: the single biggest cost saver

In a long conversation, the system prompt and early messages are identical on every turn. Without caching, you pay full price for them every single time. With prompt caching, the provider stores that stable prefix once and reuses it — charging about 10% of normal for the cached portion.

For a product like Wingman, where a single chat can run 40+ turns, this is the difference between a sustainable business and a money pit. The wingmanChat function places two cache breakpoints: one on the system prompt, one at the end of the stable history — so each new turn reuses everything before it.

🧑 In plain words
The analogy: Imagine reading a 200-page book to answer a question, then being asked a follow-up. Without caching, you re-read all 200 pages. With caching, you remember the first 199 and only read the new page. Same answer, a fraction of the effort.
⚙️ Under the hood
The mechanism: Anthropic: explicit cache_control: { type: 'ephemeral' } markers on the system block and the last stable-history message. OpenAI: automatic prefix caching (≥1024 tokens, no flag needed). Google: no explicit caching in v1. The cache lives ~5 minutes; as long as turns come within that window, every turn after the first is mostly a cache read.
Turn 1 — everything is fresh (cache write)
System
msg 1
msg 2
msg 3
cache breakpoint
new msg
💰 pay full price for system + history · write cache
Turn 2 — the prefix is already cached (cache read = ~90% off)
System
msg 1
msg 2
msg 3
msg 4
new breakpoint
new msg
💰 green = cache HIT (10% of normal price) · only the new turn is full price
05 · Resilience

Provider toggle & failover: never fully down

Three providers, each with its own account and billing. You might be testing on Google's free tier, have a few dollars of OpenAI credit, and not yet have an Anthropic API key. The system handles all of that gracefully.

Toggle: an admin sets WINGMAN_PROVIDER_CONFIG (a JSON secret) to enable/disable any provider. A provider is "available" only if enabled and its API key is set. During testing you'd leave only Google on; flip OpenAI on once funded; flip Anthropic on later.

Failover: if the chosen provider returns a quota, billing, or overload error (429, 402, 403, 5xx), the call silently retries on the next available provider — so a free tier exhausting mid-session self-heals onto the next live provider instead of breaking.

🧑 In plain words
The analogy: You have three translators on call. One's on vacation (disabled), one's free but just hit their daily limit (quota), and one is fresh and ready. The receptionist skips the first, tries the second, and when they're full, seamlessly hands your letter to the third — you never notice the switch.
⚙️ Under the hood
The logic: availableProviders(tier) builds an ordered candidate list (preferred first, then by tier preference), filtered to enabled + keyed. The failover loop tries each; on a failoverEligible error it continues, on a hard error (400 bad request) it throws. The response includes a failover trail and the provider that actually served — so you can see exactly what happened for A/B testing.
Userasks a question
wingmanChat
🔴
Anthropic
disabled / no key
skip
🟡
OpenAI
free credits
429 quota!
🟢
Google
free tier ✓
serves the reply
Usergets an answer
Failover trail (returned in the response)
failover: [
  { provider: "OpenAI", status: 429, failover_eligible: true },
]
provider: "Google"   // ← who actually served it
06 · Memory & Privacy

The data layer & privacy: what NaviStrat remembers — and protects

Everything NaviStrat knows about you lives in entities — structured database tables with built-in security. But NaviStrat isn't a generic chat app; it handles treasury-grade sensitive data: bank account numbers, trading positions, payment templates, cash forecasts, FX exposures, and client files. The data layer is designed around that reality.

What NaviStrat stores

The platform holds a wide range of entity types, each serving a specific module. Conversations and memories power Sirius. Vault documents hold reports, plans, and drafts. Bank accounts and statements (masked, never raw credentials) feed the TMS cash management module. Trading positions and investment plans drive PRISM. Payment templates and their audit logs govern the payments module. Client files and engagement updates support the consulting practice. Open items and decisions track your commitments across all modules.

Row-level security: the foundation

Every entity carries an rls block that defines who can create, read, update, and delete its records. For user-scoped data (conversations, memories, vault docs, open items, decisions), the rule is simple: read: { created_by_id: {{user.id}} } — you only ever see your own records, even though everyone's data sits in the same table. For admin-only data (GTM leads, blog management, platform activity logs, change logs), the rule is user_condition: { role: 'admin' } — only admins can access them. Some entities use compound rules: a document you own or one shared with you or an admin viewing everything.

Sensitive data categories

NaviStrat handles several categories of sensitive data, each with its own protections. Bank data — account numbers are masked, balances and transactions are stored as entities but access is controlled by BankDataAccessRule records that specify which users can see which accounts. Trading data — positions and P&L are RLS-locked to the owner. Payment data — every payment template and execution is logged in PaymentAuditLog with who/what/when. Client files — engagement notes and client data are admin-only. Contract profiles — recruiter and job application data is scoped to the owner.

Audit trails: who did what, when

For a treasury platform, auditability is non-negotiable. Three audit entities track everything: PaymentAuditLog records every payment action (created, modified, authorized, executed). PlatformActivityLog tracks user actions across all modules — which tool, what action, what entity, what result. TMSEnvSnapshot captures environment states over time so you can see what changed and when. Every entry is admin-readable and immutable.

Passcode-locked conversations

Sirius conversations can be passcode-protected for sensitive topics — M&A discussions, personal financial matters, anything you don't want visible if someone glances at your screen. The passcode is SHA-256 hashed before storage; the plaintext is never saved. A locked conversation shows a gate screen until the correct passcode is entered, and the hash is verified client-side so the server never sees the plaintext either.

Sentinel: the security & compliance module

Sentinel is NaviStrat's dedicated security layer. It maintains an access control matrix (who can do what in each module), enforces segregation of duties (the person who authorizes a payment can't be the one who executes it), tracks a risk register and fraud alerts, and maps compliance against frameworks (SOX, ISO 27001, etc.). It's the module that turns "we handle sensitive treasury data" into "we can prove we handle it right."

Multi-tenant workspace isolation

For the TMS, NaviStrat supports multi-tenant isolation through TMSWorkspace and TMSWorkspaceMember entities. Each workspace is its own isolated environment — its own bank accounts, GL accounts, counterparties, and configurations. Members are scoped to their workspace, so a user in Workspace A cannot see Workspace B's data even if they have an account on the platform.

Cross-device continuity

Conversations sync across devices in real time. Send a message on your laptop and it appears on your phone within a second — no refresh needed. The platform uses realtime subscriptions so each device is notified the moment a record changes, backed by a 60-second interval and window-focus handler for anything missed.

🧑 In plain words
The analogy: Your data is like a bank safe deposit box. Lots of boxes in the same vault, but yours has a key that only you hold. And whatever you put in from one branch instantly shows up at every other branch — because it's all one vault. For your most sensitive conversations, there's a second lock on top — a passcode only you know. And for the treasury team, there's a security guard (Sentinel) watching who opens which box and making sure no one person can both authorize and execute a withdrawal.
⚙️ Under the hood
The mechanism: Base44 entities with an rls block (read: { created_by_id: {{user.id}} } for user data; user_condition: { role: 'admin' } for admin-only). Bank data access via BankDataAccessRule. Audit via PaymentAuditLog, PlatformActivityLog, TMSEnvSnapshot. Passcodes: SHA-256 hash, client-side verification. Multi-tenant: TMSWorkspace + TMSWorkspaceMember. Cross-device sync: base44.entities.WingmanConversation.subscribe() + 60s interval + window-focus backfill. Messages capped (40 msgs, 40K chars).
What NaviStrat stores
💬
Conversations
Sirius chat history
🧩
Memories
facts · prefs · goals
📚
Vault Docs
reports · plans
🏦
Bank Accounts
masked · RLS-locked
📈
Trading Positions
P&L · exposure
💸
Payment Templates
audit-logged
📁
Client Files
engagements
✅
Open Items
tasks · decisions
all protected by
🔐
Row-Level Security
every record scoped to owner — admin entities to admin role
🔒
Passcode-Locked Chats
sensitive Sirius threads gated by SHA-256 hash
📋
Audit Trails
payment logs · activity logs · env snapshots
🛡️
Sentinel
access matrix · segregation of duties · fraud alerts
Laptopdesktop browser
Phonemobile app
Every record is scoped to its owner via row-level security — one user cannot see another's data.
07 · Reach

Integrations: talking to the outside world

The platform connects to your Google workspace through OAuth connectors — Gmail, Calendar, Drive, and Sheets. Once authorized, it can read your email, manage your calendar, save files to Drive, and sync rows to Sheets — all on your behalf.

Workflows are the automation layer: "when X happens, do Y." A new blog submission triggers a reviewer notification. A published article auto-saves to Drive. A TMS issue syncs to Sheets. They run even when no one is watching.

🧑 In plain words
The analogy: Connectors are the platform's hands and eyes — it can read your inbox, see your calendar, and file documents. Workflows are its autopilot: you set up the rule once ("when an article is published, file a copy in Drive") and it happens forever after, even while you sleep.
⚙️ Under the hood
The layer: OAuth connectors (shared mode — builder's account, all users share) with webhook support. Workflows are CNCF-style SWF definitions in base44/workflows/*.jsonc with triggers (scheduled, entity, connector webhook) and activities (invoke backend function, wait, switch). Backend functions (base44/functions/*/entry.ts) hold the actual API call logic.
🔌
OAuth Connectors
✉️
Gmail
📅
Calendar
📁
Drive
📊
Sheets
events fire
⏰
Workflows
trigger → steps → wait → branch
invoke
⚙️
Backend Functions
saveArticleToDrive · syncIssueToSheets · …
08 · Surface

The product modules

Under the hood, all these modules share the same five layers from the map above. PRISM's AI advisor and Sirius use the same wingmanChat engine. The TMS and Sentinel share the same data layer and RLS. This is why a memory you save in Sirius can inform a PRISM investment plan — they're one brain, not eight disconnected apps.

🧑 In plain words
The analogy: Think of a house with eight rooms. Each room looks different and does different things — the kitchen cooks, the office works, the gym exercises. But they share one foundation, one electrical system, and one plumbing. Fix the plumbing once, every room benefits.
📈PRISM
Treasury & trading academy, investment plans, AI advisor
🏦Ko$ha TMS
Cash management, payments, FX, debt, bank accounts
🧠Sirius / Wingman
The AI companion — chat, memory, vault, command center
📰TrueNorth
Editorial blog — treasury articles, contributors, review
🤝Nexus
Member directory, community feed, navigator profiles
🛡️Sentinel
Risk, compliance, access control, audit trails
⚙️DataForge
Master data — legal entities, GL accounts, counterparties
🎯GTM CRM
Lead pipeline, win-rate AI, engagement roadmaps
09 · Unit economics

Cost economics: credits vs direct API vs BYOK

There are three ways to pay for AI. Base44 credits are the simplest — one flat per-call price, no API keys, no billing setup. Perfect for early stage. Direct API is wholesale — you pay the provider's raw rate and unlock prompt caching, which drops long-conversation costs by ~90%. BYOK (Bring Your Own Key) is for Enterprise — the customer supplies their own provider key, so their usage never touches our margin.

🧑 In plain words
The analogy: Credits are like a hotel mini-bar — super convenient, but a $6 candy bar. Direct API is the grocery store — same candy, $1, but you have to drive there. BYOK is the enterprise cafeteria — they buy their own candy, you just provide the kitchen.
⚙️ Under the hood
The architecture: wingmanChat is the direct-API path. It holds the provider keys as secrets, runs the context router, and calls providers with caching enabled. Base44 InvokeLLM is kept only for low-volume internal tasks (chat organization, topic classification) where the flat price is actually convenient. A mode flag (context_mode: 'caching' | 'router') enables A/B testing of context-trimming quality.
Cost per 1M tokens (mid tier)
Base44 InvokeLLM (credits)
~$45
flat per-call pricing, no caching
Direct API — first turn (no cache)
$3.00
wholesale provider rate
Direct API — cached turn
~$0.40
90% of the prompt is a cache hit
BYOK (Enterprise)
cost + 0%
user supplies their own key
The crossover: Base44 credits are great for early-stage simplicity (no keys, no billing). Around 200–500 paying users, the per-call markup overtakes the cost of managing direct API keys — that's when wingmanChat takes over. Enterprise tier keeps margins safe with BYOK.
10 · In action

Examples: simple → complex

The same engine handles a one-word question and a multi-step financial analysis. The difference is which tier fires, how much context is loaded, and how many calls chain together.

🧑 In plain words
Simple: "What's SOFR?" → a quick lookup. Cheap tier, almost no context, one fast call, a tenth of a cent.
🧑 In plain words
Medium: "Draft a CFO email about our FX hedge" → the engine pulls your relevant memories (your tone, the hedge context), uses the mid tier with caching, and writes the email. About a fifth of a cent.
🧑 In plain words
Complex: "Analyze my bank statements and build a cash forecast" → this is a whole pipeline: Plaid pulls the statements, a parser extracts transactions, the full-tier LLM analyzes them, results are written to entities, and a workflow syncs a summary to Google Sheets. Five-plus calls, a few cents — and it all started from one chat message.
Simple
"What's SOFR?"→tier: cheap→Gemini Flash→1 call · ~$0.0001
Medium
"Draft a CFO email about our FX hedge"→tier: mid + 3 memories→Sonnet (cached)→1 call · ~$0.002
Complex
"Analyze my statements & build a cash forecast"→Plaid fetch → parse → LLM (full) → entity writes → workflow→Opus + Sheets sync→5+ calls · ~$0.04
11 · Putting it together

A day in the life: stick-figure edition

🟢
7:00 AMMorning greeting

You open Sirius. The daily greeting is cached (no LLM call today). It mentions last night's platform changes — surfaced from the unacknowledged changelog.

cache hit→greeting shown
🟡
9:15 AMA real question

You ask Sirius to draft a board update. Tier selector picks mid. Memories + last quarter's vault doc are loaded. wingmanChat calls Anthropic (cached) and drafts the update.

tier: mid→memories + vault→Sonnet (cached)→draft saved
🔴
11:00 AMFree tier runs dry

A quick follow-up goes to Google (cheap tier), but the free tier just hit its per-minute limit — 429. Failover silently retries on OpenAI. You never see an error.

tier: cheap→Google 429→→ failover OpenAI→reply ✓
🟣
2:00 PMHeavy analysis

You ask PRISM to analyze your portfolio and suggest a rebalance. Full tier: Opus, the full vault, Plaid account data. It writes an investment plan entity and fires a workflow to sync a summary to Sheets.

tier: full→Opus + Plaid→entity write→→ Sheets workflow
🟢
6:00 PMCross-device continuity

You close your laptop and open your phone. The afternoon's conversation is already there — realtime sync pushed it the moment each message landed. You pick up mid-thought.

realtime sync→phone caught up
End of the Sirius Document · v2

This document is itself generated and maintained as part of the platform. When the architecture changes, this page changes with it — the same engine that builds your treasury reports builds the report on itself.

← Back to NaviStrat