# Human Idea Inbox

## 2026-03-10 - Metric Stack Clarification

For this repo:
- `val_bpb` is the trunk metric, not the whole-system king metric.
- do not confuse the compression gauge with the steering wheel

Use a layered view:
- base LM: optimize held-out `val_bpb`
- runtime/controller changes: also track `total_seconds`, `num_steps`, and `peak_vram_mb`
- future selector/agent layers: target downstream utility under budget and constraints, not just language-model compression

## 2026-03-10 - Gemini "Opposite Math" Direction

Try the anti-salience direction explicitly.

Core framing:
- salience math = discrete constrained selection with decomposed, legible utility
- opposite math = unconstrained continuous synthesis under entangled black-box loss

For this repo, the practical translation is:
- remove or ablate explicit decomposed salience-style mechanisms
- favor simpler continuous end-to-end optimization over hand-crafted modular scoring terms
- test whether a plainer transformer beats the value-residual / gated additive structure

Candidate experiments:
- ablate value embeddings and their gate entirely
- ablate extra residual scalar machinery if needed
- keep the improved 5060 runtime path, then compare against the best branch on val_bpb

## 2026-03-10 - Bipolar Contextual Set Utility

Do not just mix `U` and `-U`; that collapses to a rescaled scalar.

Use separate channels:
- `P(A, x)`: promotive value
- `Q(A, x)`: aversive value
- `Ξ(A, x)`: conflict / tension, e.g. `P * Q`

For this repo, the practical translation is:
- keep promotive and aversive signals separate inside the controller
- only collapse them at the decision boundary
- use conflict to trigger caution, not to erase structure

Concrete low-risk experiment:
- promotive signal = recent loss improvement
- aversive signal = recent loss volatility / instability
- conflict = promotive * aversive
- use these to modulate LR gently on top of the current best 5060 schedule

## 2026-03-10 - Salience as Contextual Budgeted Set Selection Under Uncertain Utility

The note below supersedes the earlier rough salience draft.

I pushed it into a paper-ready shape and tightened the places where rigor usually goes feral.

Matters

* The objective is now cleanly split into latent downstream value `V*` and operational surrogate utility `Û`.
* Hard exclusions are modeled as feasibility, not as tradeable penalties.
* The propositions now make precise when itemwise ranking is exact, when redundancy gives diminishing returns, and when complementarity wrecks those guarantees.
* I included a legacy mapping appendix, but it is role-based rather than symbol-for-symbol because the exact old `S′`, `UEA`, and `SRC` formulas are not in the visible thread.

## Salience as Contextual Budgeted Set Selection Under Uncertain Utility

### Abstract

Many practical systems must choose a finite-cost subset of candidates when value depends on context, constraints, and interactions among the selected items. This paper formulates that problem as contextual budgeted set selection under uncertain utility. The true objective is latent expected downstream value, while the operational system optimizes a surrogate set utility that combines itemwise terms for novelty, goal-fit, long-run leverage, coherence, and temporal relevance with penalties for drag and redundancy and bonuses for complementarity. Hard safety, policy, validity, and compatibility constraints are represented as feasibility conditions rather than offsetting penalties. This yields a marginal-gain view of selection and clarifies the distinct roles of candidate generation, reranking, constrained optimization, control, and learning. Simpler salience-style formulas appear as restricted special cases. Stronger optimality claims require additional structure, such as modularity or submodularity, and should not be stated without those assumptions.

---

## 1. Introduction

Finite-capacity selection recurs across retrieval, recommendation, memory management, routing, and attention support. A system faces more candidates than it can carry forward, each candidate consumes some resource, and the payoff of selecting a set depends both on context and on what else is selected. Many salience-style formulations compress this into a single scalar item score. That compression is often useful as a mnemonic, but it is too coarse as a general theory.

The cleaner view is to treat salience as a constrained set-selection problem. A chooser operates under finite capacity, hard feasibility constraints, uncertain downstream utility, and interaction effects among selected items. In that setting, the right formal object is not a universal item score, but a contextual set objective together with a feasible family and a selection procedure.

This reformulation makes four clarifications. First, it separates latent downstream value from the surrogate utility the system actually estimates and optimizes. Second, it places hard constraints in feasibility rather than in soft penalties. Third, it treats redundancy and complementarity as set properties rather than pretending they are intrinsic to isolated items. Fourth, it separates scoring from optimization, control, and learning. Those separations are structural, not cosmetic.

---

## 2. Formal Setup

### 2.1 Context, candidates, and feasibility

Let `x` denote the exogenous context for a single decision episode. Context may include user state, task, session history, device, time horizon, risk posture, and any other state that changes what is valuable now.

Unless stated otherwise, `x` is held fixed while the current selection problem is solved. If selection itself changes the relevant state, the model should be applied stagewise or lifted to a sequential decision framework.

Let `I(x)` be the candidate universe available in context `x`.

Each candidate `i ∈ I(x)` has:

* hard item-level eligibility `T_i(x) ∈ {0,1}`
* cost `c_i(x) > 0` when cost-normalized scoring is used

Let `B(x)` denote the available budget or capacity.

Some hard constraints are set-level rather than item-level. Let

```text
C(A, x) ∈ {0,1}
```

denote any additional hard feasibility condition on a set `A`, such as mutual exclusion, compatibility, quota, schema, or channel constraints.

Define the eligible universe

```text
I⁺(x) = { i ∈ I(x) : T_i(x) = 1 }
```

and the feasible family

```text
F(x) = { A ⊆ I⁺(x) : Σ_{i∈A} c_i(x) ≤ B(x), C(A, x) = 1 }.
```

This formulation keeps item-level gating and set-level gating distinct while allowing both to be enforced before or during selection.

### 2.2 Latent downstream value

Let

```text
V*(A, x) = E[Y(A, x) | x]
```

denote the latent expected downstream value of selecting `A` in context `x`, where the expectation is over whatever uncertainty remains after conditioning on context.

The ideal chooser solves

```text
A*(x) ∈ argmax_{A ∈ F(x)} V*(A, x).
```

In practice `V*` is not directly known. The system therefore estimates and optimizes a surrogate objective.

### 2.3 Remarks on zero-cost and sequential cases

If genuinely zero-cost items exist, ratio scoring by `1 / c_i(x)` is not automatically meaningful. Those items should either be handled by raw marginal gain, by a separate convention, or by a richer cost model.

If adding an item changes the context materially, write states `x_0, x_1, ...` and solve a sequence of stagewise selection problems or a fuller sequential decision problem. The present formulation addresses a single decision episode with context treated as fixed during selection.

---

## 3. Feature Vocabulary and Surrogate Utility

### 3.1 Stable feature vocabulary

For each candidate `i`, define:

* `N_i(x)`: novelty or information gain
* `G_i(x)`: goal-fit or immediate relevance to the current task
* `L_i(x)`: long-run leverage, such as reuse value, retention value, compounding value, or future usefulness
* `H_i(x)`: coherence with the current trajectory, working set, or dialogue state
* `τ_i(x)`: age, staleness, or lag
* `D_i(x)`: drag, meaning friction, interruption cost, cognitive burden, latency burden, or soft risk burden

For pairs of candidates `i, j`, define:

* `K_ij(x)`: complementarity, synergy, or unlock value
* `Sim_ij(x)`: redundancy, overlap, or substitutability

Pairwise terms are taken to be symmetric unless otherwise stated. They are an approximation chosen for tractability, not a claim that all real interaction effects are pairwise.

### 3.2 Base item utility

Let `θ` collect the parameters of the surrogate model. Define the base item utility

```text
u_i(x; θ) =
  w_N(x) φ_N(N_i(x); x)
+ w_G(x) G_i(x)
+ w_L(x) L_i(x)
+ w_H(x) H_i(x)
+ w_F(x) φ_F(τ_i(x); x).
```

Here:

* `w_*(x)` are contextual weights
* `φ_N` is a learned or chosen novelty response
* `φ_F` maps age or lag to temporal relevance

This is the value of candidate `i` before interactions and soft burdens are applied.

The weights should depend on context. Research mode, emergency mode, cold start, and long-session expert use should not all share the same tradeoff profile.

### 3.3 Set utility

Define the surrogate set utility

```text
Û(A, x; θ) =
  Σ_{i∈A} u_i(x; θ)
  + Σ_{i<j, i,j∈A} K_ij(x; θ)
  - λ_div(x) Σ_{i<j, i,j∈A} Sim_ij(x; θ)
  - Σ_{i∈A} D_i(x; θ).
```

This objective adds itemwise value, rewards complementarity, penalizes redundancy, and subtracts soft burdens.

Two distinctions matter.

First, `c_i(x)` is hard resource use that appears in feasibility. `D_i(x)` is soft burden that appears in utility. They should not be conflated.

Second, `Û` is a set function for unordered selection. If order matters, the right object is a sequence function rather than a set function.

### 3.4 Operational problem

The deployed system solves

```text
Â(x) ∈ argmax_{A ∈ F(x)} Û(A, x; θ),
```

or an approximation to that problem when exact optimization is too expensive.

This makes the separation explicit:

* `V*(A, x)` is the latent objective of interest
* `Û(A, x; θ)` is the estimated operational surrogate
* `F(x)` captures hard feasibility
* the selector is the algorithm used to approximate the argmax

---

## 4. Marginal Gain and Operational Priority

The operational quantity for incremental selection is marginal gain relative to the already selected set `A`:

```text
Δ_i(A, x; θ) = Û(A ∪ {i}, x; θ) - Û(A, x; θ).
```

Under the pairwise surrogate above,

```text
Δ_i(A, x; θ) =
  u_i(x; θ)
  + Σ_{j∈A} K_ij(x; θ)
  - λ_div(x) Σ_{j∈A} Sim_ij(x; θ)
  - D_i(x; θ).
```

Define the feasible marginal-density score

```text
s_i(A, x; θ) =
  1[A ∪ {i} ∈ F(x)] · Δ_i(A, x; θ) / c_i(x).
```

This rule means:

* infeasible additions are excluded
* candidates are scored by what they add now
* normalization by cost is used only when heterogeneous costs make such normalization meaningful

If any feature depends on the evolving set rather than only on exogenous context, write that explicitly, for example `u_i(A, x; θ)`, and compute `Δ_i` directly rather than relying on the pairwise expansion above.

---

## 5. Structural Propositions

The point of this section is not to conjure theorem smoke. It is to state exactly what follows from the formulation, and no more.

### 5.1 Basic set-function notions

A set function `f : 2^V → R` is:

* **modular** if `f(A) = const + Σ_{i∈A} a_i`
* **submodular** if for all `A ⊆ B` and all `i ∉ B`,

  ```text
  f(A ∪ {i}) - f(A) ≥ f(B ∪ {i}) - f(B)
  ```
* **monotone** if `f(A) ≤ f(B)` whenever `A ⊆ B`

Submodularity is the diminishing-returns property. It matters because approximation results depend on it.

### Proposition 1. Restricted reduction to itemwise ranking

Suppose that for fixed context `x`:

1. `Û(A, x; θ) = Σ_{i∈A} v_i(x)` is modular,
2. all eligible items have equal cost,
3. feasibility is a cardinality constraint `|A| ≤ k`.

Then any set containing the `k` largest values `v_i(x)` is optimal.

**Proof sketch.** Under equal costs and a cardinality bound, feasibility depends only on the number of selected items. For any feasible set that omits a higher-valued item and includes a lower-valued one, exchanging the lower-valued item for the higher-valued one weakly increases the objective. Repeating the exchange yields a set of the top `k` values. Therefore top-`k` ranking is exact in this restricted regime.

### Proposition 2. Density ranking is not generally optimal under nonuniform costs

Even when `Û` is modular, ranking by value-per-cost is not generally optimal for 0-1 knapsack selection with nonuniform costs.

**Counterexample.** Let the budget be

```text
B = 50
```

and let three eligible items have `(value, cost)` pairs

```text
a = (60, 10)
b = (100, 20)
c = (120, 30).
```

Their value-per-cost ratios are

```text
a: 60 / 10 = 6
b: 100 / 20 = 5
c: 120 / 30 = 4.
```

A density-first rule picks `a` first, then `b`, because `a` has the highest ratio and `b` is next. The resulting set is `{a, b}` with

```text
cost = 10 + 20 = 30
value = 60 + 100 = 160.
```

Item `c` no longer fits in the remaining budget of `20`.

But the feasible set `{b, c}` has

```text
cost = 20 + 30 = 50
value = 100 + 120 = 220.
```

Since

```text
220 > 160,
```

density-first ranking is not optimal in general for 0-1 knapsack. The goblin here is discreteness: fractional-knapsack intuition does not carry over.

### Proposition 3. Redundancy-only pairwise penalties yield a submodular surrogate

Suppose `K_ij(x; θ) = 0` for all pairs and `Sim_ij(x; θ) ≥ 0`. Define

```text
b_i(x; θ) = u_i(x; θ) - D_i(x; θ)
```

and consider

```text
f(A) = Σ_{i∈A} b_i(x; θ) - λ_div(x) Σ_{i<j, i,j∈A} Sim_ij(x; θ).
```

Then `f` is submodular.

**Proof sketch.** Let `A ⊆ B` and let `p ∉ B`. The marginal gain of adding `p` to `A` is

```text
f(A ∪ {p}) - f(A) = b_p - λ_div Σ_{j∈A} Sim_pj.
```

The marginal gain of adding `p` to `B` is

```text
f(B ∪ {p}) - f(B) = b_p - λ_div Σ_{j∈B} Sim_pj.
```

Because `A ⊆ B` and all similarities are nonnegative,

```text
Σ_{j∈A} Sim_pj ≤ Σ_{j∈B} Sim_pj,
```

so

```text
f(A ∪ {p}) - f(A) ≥ f(B ∪ {p}) - f(B).
```

That is exactly diminishing returns. Therefore `f` is submodular.

### Corollary 3.1. Monotonicity requires an additional condition

Under the assumptions of Proposition 3, the surrogate is monotone if and only if all feasible marginal gains are nonnegative.

A convenient sufficient condition is

```text
b_i(x; θ) ≥ λ_div(x) Σ_{j∈A} Sim_ij(x; θ)
```

for every feasible addition of `i` to `A`.

This matters because redundancy alone gives diminishing returns, but it does not guarantee that adding an item is ever beneficial.

### Proposition 4. Positive complementarity can destroy submodularity

Suppose there exist distinct items `a` and `b` such that `K_ab(x; θ) > 0`. Then the pairwise complementarity term is supermodular, and the full surrogate may fail to be submodular.

**Proof sketch.** Consider a simplified objective with one positive pair bonus:

```text
f(A) = Σ_{i∈A} b_i + K_ab · 1[{a, b} ⊆ A].
```

Take `A = ∅`, `B = {b}`, and add item `a`. Then

```text
f({a}) - f(∅) = b_a
```

while

```text
f({a, b}) - f({b}) = b_a + K_ab.
```

Since `K_ab > 0`,

```text
b_a + K_ab > b_a.
```

So the marginal gain of adding `a` increases after `b` has already been selected. That violates diminishing returns. Hence positive complementarity can break submodularity.

### Proposition 5. Finite soft penalties cannot uniformly encode hard exclusion

Suppose ineligibility is represented by subtracting a finite penalty `M > 0` instead of using a hard gate. That is, suppose the system uses

```text
Ũ(A, x) = U(A, x) - M Σ_{i∈A} (1 - T_i(x)).
```

Then for any fixed `M`, there exist instances in which an ineligible item is still selected.

**Proof sketch.** Pick an ineligible item `q` with `T_q(x) = 0` and assume it fits within the budget. If its base contribution to utility exceeds the finite penalty, meaning

```text
Δ_q(∅, x) > M,
```

then its penalized marginal gain remains positive:

```text
Δ_q(∅, x) - M > 0.
```

So the ineligible item can still be selected if it is valuable enough. Therefore no fixed finite penalty uniformly enforces hard exclusion across instances. Hard exclusion must live in feasibility, not in a tradeable soft term.

---

## 6. Algorithmic Implications

The propositions above pin down what can and cannot be claimed.

In the modular, equal-cost, cardinality-constrained regime, top-`k` ranking is exact. That is the cleanest setting in which a single itemwise score really does solve the problem.

With modular utility and nonuniform costs, the problem becomes 0-1 knapsack. Even then, a density rule is only a heuristic, not a general optimum rule.

With redundancy-only pairwise penalties and nonnegative similarities, the surrogate becomes submodular. In that regime, diminishing returns appears naturally, and submodular-style algorithms become mathematically natural.

Once positive complementarity enters, the objective can become neither modular nor submodular. Then greedy procedures lose any blanket claim to optimality, and richer search or optimization procedures become appropriate. Depending on scale and latency budget, those may include beam search, local search, reranking loops, mixed-integer optimization, or learned approximate selectors.

The safe statement is therefore narrow:

> Algorithmic guarantees depend on the structure of the surrogate objective and the feasible family. They do not follow from a salience score alone.

---

## 7. Architectural Decomposition

A practical system should separate six layers.

### 7.1 Eligibility layer

Apply hard feasibility checks. Item-level disqualifiers go into `T_i(x)`. Set-level disqualifiers go into `C(A, x)`.

Examples include safety rules, policy rules, schema validity, source validity thresholds, compatibility checks, mutual exclusion, and quota rules.

### 7.2 Candidate generation

Recall a broad, cheap pool from the candidate universe. This stage optimizes coverage and speed, not final precision.

### 7.3 Reranker

Estimate surrogate utility terms and interactions. This stage computes itemwise value signals, drag estimates, redundancy estimates, and complementarity estimates.

### 7.4 Selector

Choose the final feasible set under `F(x)`. This stage is where greedy construction, beam search, local search, ILP, or other constrained optimization methods belong.

### 7.5 Controller

Adjust exploration, aggressiveness, budget tightening, temperature, channel multipliers, or other policy knobs that govern behavior across contexts. Exploration belongs here, not inside a single salience score.

### 7.6 Learner

Update `θ`, thresholds, transforms, interaction estimates, and controller policies from observed outcomes. Learning belongs here, not inside the same scalar equation used for online ranking.

This decomposition is useful because it prevents hard constraints, soft utility, exploration, optimization, and learning from being crammed into one overworked formula.

---

## 8. Learning and Evaluation

The separation between `V*` and `Û` has direct learning consequences.

Observed feedback typically arrives only for the selected set, not for all counterfactual feasible sets. That means learning the surrogate is partly a counterfactual estimation problem. A system may estimate component models for itemwise terms and pairwise terms, fit a direct predictor of set-level outcomes, or combine both approaches.

A generic learning target is to estimate parameters `θ` so that higher surrogate utility corresponds to better observed outcomes under deployment-relevant conditions. In symbolic form, one may fit `θ` by minimizing a loss of the form

```text
Σ_t ℓ( ψ(A_t, x_t; θ), y_t ) + Ω(θ),
```

where `A_t` is the selected set in episode `t`, `y_t` is the observed outcome, `ψ` is the model output used for training, and `Ω` is any regularizer or structural penalty.

Two evaluation cautions follow.

First, when interactions matter, itemwise ranking metrics can become a misleading proxy for decision quality. A system that ranks individual items well can still assemble poor sets.

Second, evaluation should track the level at which the system acts. If the selector chooses sets, then at least part of the evaluation should be set-level or policy-level rather than purely item-level.

This is one more reason the latent-versus-surrogate split is not just notation cleanup. It affects the learning target and the evaluation protocol.

---

## 9. Domain Instantiations

### 9.1 Retrieval and RAG

In retrieval settings:

* `G` captures query relevance
* `L` captures downstream usefulness for reasoning or future reuse
* `H` captures coherence with the current dialogue or working set
* `τ` captures source age or lag
* `Sim` captures chunk overlap
* `K` captures multi-hop complementarity across evidence pieces
* `D` captures token cost, latency, or distraction burden
* `T` and `C` capture trust, source validity, safety, and policy constraints

### 9.2 Training data selection

In training data selection:

* `N` captures undercoverage or surprise
* `G` captures fit to the current training objective
* `L` captures long-run capability gain
* `H` captures curriculum coherence
* `τ` captures distribution lag where relevant
* `Sim` captures duplication or oversampling overlap
* `K` captures complementary samples that unlock generalization
* `D` captures labeling cost, noise burden, or instability risk
* `T` and `C` capture license, validity, and policy constraints

### 9.3 Product and engagement systems

In product flows:

* `G` captures immediate task completion value
* `L` captures retention or future utility
* `H` captures journey coherence
* `τ` captures timeliness
* `Sim` captures recommendation redundancy
* `K` captures next-step unlocks or useful bundles
* `D` captures friction and interruption burden
* `T` and `C` capture trust, safety, and experience constraints

### 9.4 Human operator mode

For human reasoning, a compact salience heuristic can still be useful. But it should be treated as an operator shorthand over this substrate, not as the substrate itself.

---

## 10. Scope of Claims

The model supports strong structural claims and weaker universal claims. Keeping those separate prevents puffed-up nonsense.

Safe claims:

* Many practical selection systems can be modeled as contextual feasible-set selection under uncertain utility.
* Itemwise salience formulas are restricted special cases of that broader problem.
* Hard exclusions are more robustly modeled as feasibility conditions than as finite penalties.
* Redundancy can produce diminishing returns, while complementarity can destroy them.

Unsafe claims unless extra assumptions are stated:

* One scalar item score explains all selection behavior.
* Greedy selection is optimal in general.
* Value-per-cost ranking solves the nonuniform-cost case.
* Hard constraints can always be replaced by sufficiently large soft penalties.

---

## 11. Conclusion

The useful general principle is not that one salience equation explains everything. The useful principle is that finite-capacity selection is a contextual feasible-set problem under uncertain utility. Candidates have costs. Some candidates are ineligible. Value depends on context. Redundancy and complementarity are properties of sets, not just of isolated items. The deployed chooser acts on an estimated surrogate through feasible marginal gains, often normalized by cost when that normalization is meaningful.

That framing preserves what was good in the original salience idea while removing notation drift, layer confusion, and fake universality.

---

## Appendix A. Compact Canonical Form

The full problem can be summarized as follows.

Given context `x`, define a feasible family

```text
F(x) = { A ⊆ I⁺(x) : Σ_{i∈A} c_i(x) ≤ B(x), C(A, x) = 1 }.
```

Let latent downstream value be

```text
V*(A, x) = E[Y(A, x) | x].
```

Let surrogate utility be

```text
Û(A, x; θ) =
  Σ_{i∈A} u_i(x; θ)
  + Σ_{i<j, i,j∈A} K_ij(x; θ)
  - λ_div(x) Σ_{i<j, i,j∈A} Sim_ij(x; θ)
  - Σ_{i∈A} D_i(x; θ).
```

Then the ideal and operational problems are

```text
A*(x) ∈ argmax_{A ∈ F(x)} V*(A, x)
```

and

```text
Â(x) ∈ argmax_{A ∈ F(x)} Û(A, x; θ).
```

The incremental priority of candidate `i` relative to current set `A` is

```text
Δ_i(A, x; θ) = Û(A ∪ {i}, x; θ) - Û(A, x; θ)
```

and, when meaningful,

```text
s_i(A, x; θ) =
  1[A ∪ {i} ∈ F(x)] · Δ_i(A, x; θ) / c_i(x).
```

---

## Appendix B. Legacy Mapping of `S′`, `UEA`, and `SRC`

Because the exact legacy formulas are not reproduced in the visible thread, the mapping below is structural rather than symbol-for-symbol. That is still enough to place each old object in the new framework without pretending to recover algebra that is not on the page.

### B.1 Legacy scalar salience score `S′`

If `S′` was an itemwise scalar used to rank candidates, it maps to one of two places.

If the old score represented raw item value, it maps most closely to

```text
u_i(x; θ) - D_i(x; θ)
```

under the simplifying assumptions that interactions are ignored and feasibility is itemwise.

If the old score represented value relative to resource use, it maps more closely to the marginal-density score

```text
s_i(A, x; θ).
```

Under the most restrictive special case,

```text
S′_i(x) ≈ [u_i(x; θ) - D_i(x; θ)] / c_i(x)
```

when:

* `A = ∅`
* `K_ij = 0`
* `Sim_ij = 0`
* `T_i(x) = 1`
* set-level feasibility is absent or trivial

The key upgrade is that the new framework replaces a fixed scalar `S′_i(x)` with a stateful marginal quantity `s_i(A, x; θ)` once interactions matter.

### B.2 Legacy expected-utility object `UEA`

If `UEA` was intended to denote true downstream performance, then it belongs with the latent objective

```text
V*(A, x).
```

If `UEA` was used as an operational scoring function inside the deployed system, then it belongs with the surrogate objective

```text
Û(A, x; θ)
```

or with one of its components.

Those two roles should not share one symbol. One is the unknown target of interest. The other is the model the system can actually compute.

### B.3 Legacy controller or selector object `SRC`

If `SRC` governed exploration, aggressiveness, routing, or channel choice, it belongs in the controller layer.

If `SRC` carried out the constrained choice of a feasible set, it belongs in the selector layer.

If `SRC` was written as though it were another additive term in the score, the new framework recommends splitting it apart. Controllers and selectors act on the utility model and feasibility conditions. They are not themselves utility terms.

### B.4 Legacy trust or safety modifiers

Any binary disqualifier that previously appeared as a multiplier or penalty should move into either

```text
T_i(x)
```

or

```text
C(A, x).
```

Any graded source-quality or confidence effect that is not a hard exclusion can remain as a soft term, but it should be represented explicitly as such rather than smuggled in as pseudo-eligibility.

### B.5 Legacy freshness and novelty multipliers

Fixed freshness or novelty multipliers map naturally to either contextual weights

```text
w_N(x), w_F(x)
```

or response functions

```text
φ_N(·; x), φ_F(·; x).
```

This is cleaner because it distinguishes “how much novelty matters now” from “how the system responds to a given novelty level.”

---

## Appendix C. Drop-in Claim Language

These are paper-safe claims that can be inserted into an introduction or discussion without overpromising.

1. **Structural framing.**
   Many retrieval, recommendation, routing, and memory systems can be modeled as contextual feasible-set selection under uncertain utility.

2. **Operational framing.**
   A useful operational strategy in such systems is to estimate feasible marginal utility, often normalized by cost, using a surrogate that includes drag, redundancy, and complementarity.

3. **Special-case framing.**
   Itemwise salience scores are exact only in restricted regimes where the objective is effectively modular and the feasible family is simple enough to collapse set selection into ranking.

4. **Constraint framing.**
   Hard policy, safety, and validity constraints are better represented as feasibility conditions than as tradeable penalties.

5. **Guarantee framing.**
   Approximation or optimality claims require additional assumptions about the surrogate objective and the feasible family, such as modularity, monotonicity, or submodularity.

At this point the manuscript has a clean spine: abstract, formal model, propositions, algorithmic implications, architecture, learning remarks, and legacy mapping. The next natural form is either straight LaTeX or a shorter essay version aimed at a broader technical audience.

Paste rough ideas here while the agent is running.

No schema is required. Dump math, pseudocode, paper pointers, constraints, half-baked hunches, or direct commands.

Useful optional conventions:
- Put newer ideas near the top.
- Prefix a line with `!` if you want it tried soon.
- Prefix a line with `constraint:` if it should be treated as a hard instruction, not a soft suggestion.

The agent should reread this file from disk before every experiment instead of relying on its initial context.
