Files
ruvnet--RuView/harness/wifi-densepose-sar/src/router.ts
T
ruv 1b220c8d53 feat(harness): scaffold wifi-densepose-sar-harness with darwin, router, flywheel
Mints a real MetaHarness (via vendor/metaharness's published `npx
metaharness analyze --scaffold`, template vertical:coding, host
claude-code) for the wifi-densepose-sar crate: architect/implementer/
reviewer/test-writer agents, doctor/review-diff commands, MCP server,
Claude Code plugin -- following the same pattern as harness/ruview/
(ADR-182) and harness/homecore/ (ADR-285).

Adds real wiring for the three pieces this was scoped around:
- Darwin Mode (@metaharness/darwin) -- wired by the scaffold itself
  (npm run evolve / evolve:dry).
- Router (@metaharness/router) -- src/router.ts, a real k-NN
  cost-optimal Router over two example model tiers. Its labelled
  examples are illustrative seed data (see the file's honesty note),
  not measured eval-log observations; the routing mechanism itself is
  real and tested.
- Flywheel (@metaharness/flywheel) -- src/flywheel.ts, the real
  propose/evaluate/gate/promote loop wired with a SYNTHETIC proposer
  and evaluator (dataSource: 'SYNTHETIC', no model call). Proves the
  wiring end-to-end: a real signed, independently-replayable lineage,
  promoting each generation once the evaluator's noopRate actually
  moves (the default gate requires it to strictly improve -- a
  constant noopRate, even a "good" one, blocks every promotion
  forever, which the first version of this evaluator hit and the
  final version fixes).

14/14 tests pass (5 router + 5 flywheel + 4 install-smoke), `npm run
build` clean under strict TypeScript, CLI commands (route, flywheel)
verified manually. `.harness/manifest.json` is stale relative to the
router/flywheel additions -- this scaffold has no manifest:update
script (unlike harness/homecore/); documented as a known gap in the
harness's own README.
2026-07-31 00:36:38 -04:00

78 lines
3.5 KiB
TypeScript

// SPDX-License-Identifier: MIT
//
// Cost-optimal task routing for the wifi-densepose-sar harness, via
// @metaharness/router (ADR-040's DRACO Phase-2 finding, productized): route
// each agent query to the cheapest model predicted to clear a quality bar,
// instead of sending every query to the frontier tier by default.
//
// HONESTY NOTE: the candidate `examples` below are SEED/ILLUSTRATIVE data —
// four hand-picked (embedding, quality) points per candidate, not measured
// eval-log observations. They exist so `sarTaskRouter` is a real, runnable
// k-NN router out of the box (see __tests__/router.test.ts), not so its
// routing decisions should be trusted for production cost savings. Replace
// `SAR_ROUTER_CANDIDATES[*].examples` with real (query embedding → quality
// achieved) rows from your own eval logs before relying on this for
// production routing — see @metaharness/router's README ("the more
// examples, the closer it gets to the per-query oracle").
import { Router, type RouterCandidate } from '@metaharness/router';
/**
* A 4-axis feature embedding for a harness query, used only to pick a
* nearby labelled example — NOT a real text embedding. Axes (each 0..1):
* [0] physicsExplanation — "explain the range-resolution formula"-shaped
* [1] codeReview — "review this diff for correctness"-shaped
* [2] numericalDebugging — "why did this reconstruction test fail"-shaped
* [3] docWriting — "write/update the tutorial"-shaped
* A caller with a real embedding model should project onto whatever
* dimensionality that model produces instead — the router only needs
* consistent vectors, not these specific four axes.
*/
export type SarTaskEmbedding = readonly [number, number, number, number];
export const SAR_ROUTER_CANDIDATES: RouterCandidate[] = [
{
id: 'cheap-tier',
costPerMTok: 1,
examples: [
{ embedding: [1, 0, 0, 0], quality: 0.9 }, // physics explanations: cheap tier does fine
{ embedding: [0, 0, 0, 1], quality: 0.85 }, // doc writing: cheap tier does fine
{ embedding: [0, 1, 0, 0], quality: 0.55 }, // code review: cheap tier is weak
{ embedding: [0, 0, 1, 0], quality: 0.5 }, // numerical debugging: cheap tier is weak
],
},
{
id: 'frontier-tier',
costPerMTok: 15,
examples: [
{ embedding: [1, 0, 0, 0], quality: 0.95 },
{ embedding: [0, 0, 0, 1], quality: 0.93 },
{ embedding: [0, 1, 0, 0], quality: 0.92 }, // code review: frontier tier needed
{ embedding: [0, 0, 1, 0], quality: 0.9 }, // numerical debugging: frontier tier needed
],
},
];
/**
* Cost-optimal router for the harness's four query shapes above. `qualityBar`
* of 0.8 matches the router README's worked example: return the cheapest
* candidate predicted to clear 80% quality, or the best-predicted candidate
* if none do.
*/
export const sarTaskRouter = new Router({
qualityBar: 0.8,
candidates: SAR_ROUTER_CANDIDATES,
// k=1: each candidate has only 4 (orthogonal, one-hot) examples covering
// the 4 task axes. The router's default k=5 would average ALL of a
// candidate's examples regardless of query similarity once a candidate
// has <=5 examples, collapsing every query to the same prediction. k=1
// makes it pick the single nearest labelled task type, which is what
// this small illustrative dataset is shaped for.
k: 1,
});
/** Route one query embedding to the cost-optimal model tier. */
export function routeSarQuery(queryEmbedding: SarTaskEmbedding) {
return sarTaskRouter.route([...queryEmbedding]);
}