# Llama MCP Bridge

Full-duplex MCP/ACP control over llama.cpp. Every single byte that passes through the model is inspectable and modifiable.

## What This Gives You

### Read Access (Inspect Everything)
- **Logits** - Full vocab distribution at every position
- **Embeddings** - Token embeddings, position embeddings, layer outputs
- **KV Cache** - Every key/value vector in context memory
- **Attention** - Q/K/V vectors, attention weights, attention output
- **Activations** - Hidden states at every layer
- **Weights** - Model parameters (wq, wk, wv, wo, ffn gates, etc.)

### Write Access (Modify Everything)
- **Inject tokens** at any position
- **Inject embeddings** - bypass tokenization entirely
- **Modify logits** - force specific tokens or suppress others
- **Edit KV cache** - insert, delete, or alter "memories"
- **Steer activations** - modify hidden states at any layer
- **Change weights** - permanently alter model behavior

### Feature Steering (Golden Gate Claude Style)
- Load SAE (Sparse Autoencoder) weights
- Identify interpretable features
- Steer features to change model behavior
- Works with Gemma Scope SAEs

## Build

```bash
# Build llama.cpp with the bridge
cd llama-mcp-bridge
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release

# Build MCP server
cd mcp-server
npm install
npm run build
```

## Usage

### As MCP Server

Add to your MCP client config:
```json
{
  "mcpServers": {
    "llama": {
      "command": "node",
      "args": ["C:\\custollama\\llama-mcp-bridge\\mcp-server\\dist\\index.js"]
    }
  }
}
```

### Example: Full Control

```javascript
// Load model
await llama_load_model({ model_path: "models/gemma-2-9b-q4_k_m.gguf", n_ctx: 8192 });

// Tokenize
const { tokens } = await llama_tokenize({ text: "Hello world" });

// Run forward pass
const { logits } = await llama_decode({ tokens });

// INSPECT: Get logits for position 0, see full vocab distribution
const allLogits = await llama_get_logits();
console.log("Top tokens:", getTopK(allLogits, 10));

// MODIFY: Force the model to output "Golden Gate Bridge"
await llama_set_logits({ logits: boostToken(logits, "Golden") });

// Sample next token
const nextToken = await llama_sample({ logits: modifiedLogits });

// INJECT: Skip tokenization, inject raw embedding
await llama_inject_embedding({ 
  embedding: customEmbedding,  // Your custom vector
  position: 5 
});

// INSPECT KV CACHE: See what model "remembers"
const kv = await llama_get_kv_cache({ layer: 0, pos: 0 });
console.log("Key shape:", kv.k.length);
console.log("Value shape:", kv.v.length);

// MODIFY KV CACHE: Inject a "memory"
await llama_set_kv_cache({
  layer: 12,
  pos: 100,
  k: customKeyVector,
  v: customValueVector
});

// STEER FEATURE (Golden Gate Claude style)
await llama_steer_feature({ 
  feature: 3439,  // "Golden Gate Bridge" feature from SAE
  strength: 5.0,
  layer: 12 
});

// Continue generation with steering active
const output = await llama_generate({ prompt: "What should I visit?", max_tokens: 100 });
// Model will be obsessed with Golden Gate Bridge!
```

## Hook Points

Register hooks at any point in the forward pass:

- `tokenize_pre` / `tokenize_post`
- `embed_pre` / `embed_post`  
- `attention_pre` / `attention_post`
- `ffn_pre` / `ffn_post`
- `layer_pre` / `layer_post`
- `logits_pre` / `logits_post`
- `sample_pre` / `sample_post`
- `kv_cache_read` / `kv_cache_write`

## Architecture

```
┌─────────────────────────────────────────────────────────────────┐
│                        MCP Server (TypeScript)                   │
│  Exposes 40+ tools for complete control over inference          │
└───────────────────────────┬─────────────────────────────────────┘
                            │ FFI (koffi)
┌───────────────────────────▼─────────────────────────────────────┐
│                   Llama MCP Bridge (C++)                         │
│  Hooks into every layer, captures all tensors, enables steering │
└───────────────────────────┬─────────────────────────────────────┘
                            │
┌───────────────────────────▼─────────────────────────────────────┐
│                      llama.cpp                                   │
│  GGUF inference engine with full access to internals             │
└─────────────────────────────────────────────────────────────────┘
```

## Papers This Enables Research On

- [Scaling Monosemanticity (Golden Gate Claude)](https://transformer-circuits.pub/2024/scaling-monosemanticity/) - Feature steering
- [Gemma Scope](https://arxiv.org/abs/2408.05147) - SAE interpretability
- Activation engineering / representation engineering
- Mechanistic interpretability
- Model editing / weight surgery

## Model Recommendations

Works with any GGUF model. For interpretability research:
- **Gemma 2 9B** - Has Gemma Scope SAEs available
- **Llama 3.1 8B** - Good general purpose model
- **Phi-3 Mini** - Smaller for faster experimentation
