MONIKA / Meta-Orchestrated Novelty & Invariant Kernel Architecture

Exploring salience-driven compute allocation in language model architectures

MONIKA (Meta-Orchestrated Novelty & Invariant Kernel Architecture) is the codename for the SalienceOS Seed runtime implemented in this repository. It wires together `SalienceRuntime` (`runtime/orchestrator.py`), the salience sensor stack (`core/sensors/`), controller policy (`core/controller/policy.py`), state-space backbone (`core/operators/sass.py`), and adaptive memory/tooling subsystems to study salience-governed compute. Instead of running a fixed stack of blocks, the runtime measures state on every step, scores actions with the S′ controller, and only executes the operators that are expected to deliver value.

Motivation

Standard transformer architectures execute the same operations regardless of input complexity—every token attends to every other token at every layer. This works well but treats trivial and complex inputs identically. We're exploring whether explicitly modeling computational value (salience) and selecting operators accordingly offers benefits for efficiency or capability.

Core Mechanism

On each step MONIKA derives an eight-channel salience vector (uncertainty, novelty, alignment, progress, cost, drag, truth, coherence). `SalienceControllerPolicy` blends S′ scores (`core/controller/s_prime.py`) with compute-auction bids and bandit priors before handing a decision to `SalienceRuntime`. The event-driven scheduler (`core/scheduler/`) enforces cooldowns and budget limits, while the runtime conditionally runs SASS, verification, reflection, memory verbs, or tool hooks.

S' = (w₁·ΔA + w₂·R + w₃·M) × C × e-λt × (1 - kφ)
ΔA Novelty
R Retention
M Payoff
C Continuity
φ Fatigue
w₁,w₂,w₃ Learned weights
# Simplified runtime loop def run_step(state): # Measure current state salience = sensor_bank.tick(state, memory, meta) # Score all possible actions decision = controller.choose(salience, meta, budget) # Execute chosen operator if decision.operator == SASS: output = sass_core.forward(hidden_states) elif decision.operator == VERIFY: result = verifier.run(state) elif decision.operator == REFLECT: reflection_step(decision, state) # Update weights based on outcome bandit_trainer.update(decision.action, reward)

Implementation Components

Sensor Bank

Eight MAD-normalised sensors emit the canonical salience vector (uncertainty, novelty, alignment, progress, cost, drag, truth, coherence) using `MedianMADNormalizer` and dedicated sensor classes.

core/sensors/bank.py

Controller Policy

`SalienceControllerPolicy` applies S′ scoring (`core/controller/s_prime.py`), compute auction bids, hysteresis, cooldowns, and a bandit prior store to pick the next operator/patch/depth tuple.

core/controller/policy.py

SASSCore

Stacked state-space blocks with gated convolutions, RoPE rotation, and optional hyper-adapter deltas provide linear-time token processing when selected.

core/operators/sass.py

Runtime Orchestrator

`SalienceRuntime` stitches together sensors, controller, scheduler, compute auction, verifier, memory operator, graph reasoner, and teleporter while tracking budgets, salience histories, and metrics.

runtime/orchestrator.py

Sparse Teleporter

`SparseJumpTeleporter` keeps a per-sequence KV cache and injects residuals on salience-triggered hops, approximating selective long-range recall without full attention cost.

core/operators/sparse_jump.py

ProtoLanguageModel

Lightweight autoregressive learner sharing SASSCore, featuring dynamic vocabulary growth, age-weighted loss, module plug-ins, EMA parameters, and evaluation helpers.

proto_lm/trainer.py

Implementation Status

What's Implemented

✓ Salience sensor bank
✓ S' scoring function
✓ Bandit controller
✓ Runtime orchestrator
✓ State-space core
✓ Structured memory
✓ Adaptive coordination
✓ Synthetic baseline evaluation

The core runtime loop is functional. `eval.synthetic_baseline.py` measures lightweight metrics on the compact synthetic splits, and the telemetry bus (`telemetry.py`) streams parameter updates, runtime metrics, and ingestion events.

Validation & Tooling

Integration tests and CLI tooling exercise the runtime, adaptive coordinator, and evaluation harness:

# Eval synthetic baseline checkpoint python eval.synthetic_baseline.py --checkpoint storage/proto_lm/synthetic_baseline.pt # Runtime dashboard (event-driven scheduler + telemetry) python -m salience_os_seed.runtime.ui.cli --generator spiky # Integration tests (runtime, adaptive coordinator, conversation) pytest tests/test_adaptive_integration.py tests/test_controller.py tests/test_conversation.py -q

These scripts exercise the same modules used in production runs—covering controller hysteresis, adaptive gating, structured memory persistence, and the synthetic evaluation pipeline.

Running the System

# Run conversational interface python -m salience_os_seed.conversation.cli # Evaluate checkpoint python eval.synthetic_baseline.py --checkpoint storage/proto_lm/synthetic_baseline.pt # Launch dashboard UI python -m salience_os_seed.runtime.ui.cli --generator baseline

Architectural Comparison

Standard Transformers

  • Fixed layer stack executes on every input
  • O(n²) attention always computed
  • Single operation: multi-head attention
  • No explicit resource allocation
  • Weights frozen after training

MONIKA Architecture

  • Dynamic routing based on salience invariant
  • Can skip expensive ops if low-value
  • Multiple operators: process, verify, reflect
  • Explicit salience measurement
  • Online bandit learning

The key difference: transformers execute uniformly, whereas MONIKA measures first then decides. Whether that measurement overhead is worth the potential allocation savings is an empirical question we're investigating.

Limitations & Open Questions

Known Tradeoffs

Complexity
More components than standard architectures—harder to debug, tune, and maintain
Overhead
Sensor evaluation and scoring add per-step cost; only beneficial if allocation savings exceed measurement cost
Scale
Currently tested on small models (synthetic baseline scale); behavior at larger scale unknown
Benchmarking
Need comprehensive comparisons against baseline transformers across diverse tasks
Learning
Bandit learning may be slower to converge than end-to-end gradient descent in some scenarios

Approach

MONIKA explores whether treating compute allocation as an explicit optimization problem (via salience measurement) offers advantages over implicit allocation (via attention). The S' equation—the salience invariant—is one instantiation of this idea, representing a kernel of adaptive resource management.

The implementation prioritizes:

  • Modularity: Each component (sensors, controller, operators) is independently testable
  • Measurability: Telemetry exposes salience values, decisions, and outcomes for analysis
  • Adaptability: Bandit learning allows runtime adjustment without full retraining
  • Transparency: Decisions are interpretable through salience scores and controller logs

Whether MONIKA offers practical advantages over well-optimized transformers remains an open empirical question. The goal is to demonstrate feasibility and gather data on where salience-driven allocation helps or hurts.

System Architecture Diagram

INPUT Sensor Bank 8 sensors measure state • Uncertainty • Novelty • Progress • Truth • Drag • Cost • Coherence S' Scoring Engine w₁·ΔA + w₂·R + w₃·M × C × (1-kφ) Computes salience for all actions Controller Policy Bandit learning + hysteresis Chooses highest-value operator SASS State-space processing VERIFY Check correctness REFLECT Introspect patterns MEMORY Query/update structures TOOLS External resources Execute & Observe Record outcome, compute reward Adaptive Learning Update bandit weights Adjust truth guards Tune elegance thresholds OUTPUT Response + Metrics Gated by truth/elegance Raw text or conversation Core Innovation: Compute WHAT to compute based on expected value

Legend:

🟡 Yellow pulse = Data in motion

🟣 Purple = Decision & control

🟢 Green = Execution operators

🔵 Blue = Adaptive learning

- - - Dashed = Feedback loop

Interactive Flow: Watch the glowing particle travel through MONIKA. Data flows top-to-bottom through measurement → scoring → selection → execution → learning. The yellow feedback loop shows how adaptive learning continuously improves future decisions. Hover over components to see them respond!

This is an experimental research system. Code, documentation, and evaluation scripts are available for review. Feedback and empirical comparisons welcome.