mirror of
https://github.com/ruvnet/RuView
synced 2026-07-24 17:43:20 +00:00
e6f26e9ac9
* docs(adr): deep review of the RuView npm surface — ADR-263/264/265 optimization strategies
ADR-263 — @ruvnet/ruview@0.1.0 harness review (O1–O9):
- HIGH: claim-check CLI fails open on empty input (no --text/--file -> PASS exit 0)
- HIGH: MCP stdio server head-of-line blocking (spawnSync verify/calibrate up to 600s)
- MEASURED: optionalDependencies triple the cold npx install (4 pkgs/620kB/71 files
vs 1 pkg/172kB/22 files with --omit=optional) for a path that never imports them
- maxBuffer truncation, python -c port interpolation, version drift, duplicate skills,
guardrail METRIC_TERMS substring false positives ('map'/'F1' — found by dogfooding
claim-check on these very ADRs), zero CI
ADR-264 — @ruvnet/rvagent@0.1.0 + @ruv/ruview-cli review (O1–O9), verified against
the published registry tarball:
- HIGH: exports.require -> dist/index.cjs which is never built nor published
- MEASURED: 44 dead source-map files = 62,698B of the 188kB unpacked payload
- stdio-only server described as dual-transport; mixed dot/underscore tool names;
double Zod validation + hand-duplicated advertised schemas; 2-fd leak per training
job; unbounded body in the unwired HTTP scaffold; dead detectCogBinary candidates;
ruview bin-name collision
ADR-265 — cross-cutting npm distribution strategy: npm-packages.yml CI matrix
(test + pack-content/size gate + tarball-install smoke test), publish-from-CI-only
with npm provenance, version single-sourcing from package.json, bin/namespace
ownership (ruview bin belongs to @ruvnet/ruview), claim-check on package READMEs.
Docs only — no runtime code changed. Index/CHANGELOG/CLAUDE.md/README counts updated.
Co-Authored-By: claude-flow <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01WrGfTGKv1oWZ6iwXZACULz
* fix(npm): implement ADR-263/264/265 — harness fail-closed + async MCP, rvagent packaging/transport/naming, npm CI+provenance gate
ADR-263 (@ruvnet/ruview 0.2.0), O1-O9:
- claim-check fails closed on empty input (CLI exit 2, empty_text tool error)
- MCP stdio server dispatches tools/call asynchronously (promise-based spawn);
ping answers while a 3s fake verify runs — pinned by new e2e test
- optionalDependencies dropped: cold npx installs exactly 1 package
(MEASURED: was 4 pkgs/620kB/71 files via npm i in a clean prefix)
- bounded rolling output tails replace spawnSync 1MiB maxBuffer
- node_monitor port passed via sys.argv, never spliced into python -c source
- serverInfo.version read from package.json; resources/prompts stubs
- skills single-sourced: prepack sync script generates .claude/skills/ copies
- which() = memoized dep-free PATH scan
- tools underscore-canonical (ruview_claim_check, ...) + dotted aliases
- guardrail precision: word-boundary map/f1/auc/iou, code-span + F1/O2 label
scrubbing, quantitative-claims-only; packaging reproducer hints
- 30/30 tests (was 17), incl. concurrency e2e + fail-open regression pins
ADR-264 (@ruvnet/rvagent 0.2.0), O1-O9:
- exports fixed: types-first, phantom dist/index.cjs require target removed
- tarball map-free: 127,704B unpacked / 46 files / 0 maps (MEASURED,
npm pack --dry-run; was 188kB incl. 44 maps referencing unshipped src)
- Streamable HTTP actually wired behind RVAGENT_HTTP_PORT: one transport +
one MCP server per session (mcp-session-id routing), 1MiB body cap (413),
port-aware localhost origin gate; dual-transport description now true
- tools renamed underscore-canonical with dotted router-only aliases
- single Zod validation gate; advertised inputSchema generated from the same
Zod source (zod-to-json-schema)
- train_count: parent log fds closed (was leaking 2/job); job records
persisted to <jobsDir>/<id>.json (job_status survives restarts); bounded
log-tail reads
- detectCogBinary probes its candidates instead of dead-coding them
- version from package.json; @types/express dropped; @types/jest -> 29
- README rewritten to match reality (no phantom subcommands/policy layer)
- 99/99 jest tests (incl. new session/body-cap suite + previously-broken
manifest suite); stdio handshake + HTTP session flow smoke-tested live
ADR-265 D1-D4:
- .github/workflows/npm-packages.yml: 3-package x Node 20/22 gate — tests,
version-literal grep (D3), pack-content/size gate, tarball-install smoke
test (catches the ADR-264 F1 class), README claim-check (D4)
- .github/workflows/ruview-npm-release.yml: publish from CI only with
npm publish --provenance
- @ruv/ruview-cli bin renamed ruview-cli (ruview bin belongs to
@ruvnet/ruview); version single-sourced
- ci.yml NODE_VERSION 18 -> 20
ADR statuses updated to Accepted/implemented; harness manifest re-pinned;
ADR-263/264/265 + both package READMEs pass claim-check.
Co-Authored-By: claude-flow <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01WrGfTGKv1oWZ6iwXZACULz
* perf(rvagent): lazy-load HTTP transport + memoize generated tool schemas
stdio time-to-first-response ~242ms -> ~189ms (-22%; MEASURED, median of
repeated initialize round-trips against dist/index.js in this container).
- ./http-transport.js now imported lazily inside the RVAGENT_HTTP_PORT
branch: it chain-loads the MCP SDK streamableHttp module (~48ms MEASURED
via per-module import() timing) which the default stdio path never uses
- toolInputJsonSchema memoized per tool: schemas are static for the process
lifetime; under the session-per-server HTTP model every session calls
tools/list, so stop re-walking the Zod tree each time
No behavior change: 99/99 jest tests; HTTP session flow re-smoke-tested
through the lazy import path (initialize -> 200 + mcp-session-id).
Profiled @ruvnet/ruview too and left it alone: 50ms CLI startup vs ~29ms
bare 'node -e ""' floor on the same box (MEASURED) — already near the
interpreter floor with zero dependencies.
Co-Authored-By: claude-flow <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01WrGfTGKv1oWZ6iwXZACULz
* ci(ruview-cli): pass jest --passWithNoTests so the private no-test package doesn't fail the npm-packages matrix
Co-Authored-By: claude-flow <ruv@ruv.net>
* fix(npm): address 10 verified review findings in harness + rvagent before 0.2.0 publish
harness/ruview (@ruvnet/ruview):
- guardrails: digit gate now sees numbers inside code spans; F1-style
metric tokens followed by ':' or a nearby number are no longer scrubbed
(fail-open regressions in the honesty gate)
- mcp-server: tools/call requests serialize through a FIFO promise chain
(hardware/mutating tools never overlap) while ping/tools/list stay
immediate; stdin close drains in-flight responses before exit
- tools: which() no longer memoizes negative lookups
tools/ruview-mcp (@ruvnet/rvagent):
- index: realpath invoked-directly guard — library import no longer
connects a stdio transport to the consumer's process
- http-transport: explicit allowedOrigins is exact-match only (localhost
any-port convenience applies only with no configured allowlist);
session map gains maxSessions=64 + 5min idle TTL sweep
- train-count: job records persist the child pid and reconcile stale
'running' status after a server restart (exit-code marker or dead pid)
- config: cog binary candidates ordered by process.arch
.github/workflows/ruview-npm-release.yml: port the full ADR-265 D1 gate
(version-literal check, unpacked-size budget, tarball-install smoke test)
from npm-packages.yml so the publish path enforces what the header claims.
Tests: harness 30→36, rvagent 99→112, all passing.
Co-Authored-By: claude-flow <ruv@ruv.net>
---------
Co-authored-by: Claude <noreply@anthropic.com>
149 lines
6.9 KiB
JavaScript
149 lines
6.9 KiB
JavaScript
// SPDX-License-Identifier: MIT
|
||
// RuView harness guardrails — the "prove everything" rule made executable.
|
||
//
|
||
// The project was accused of AI-slop; the cultural fix is that every accuracy
|
||
// number must be tagged MEASURED (with a reproducer) or CLAIMED/SYNTHETIC, and
|
||
// the retracted "100% accuracy" framing must never reappear untagged. This module
|
||
// is the static enforcement of that, shared by the `ruview_claim_check` MCP tool,
|
||
// the `npx ruview claim-check` CLI, and the claude-code pre-output hook.
|
||
|
||
/** Phrases that signal a quantitative accuracy claim (safe as substrings). */
|
||
const METRIC_TERMS = [
|
||
'accuracy', 'pck', 'precision', 'recall',
|
||
'mpjpe', 'error rate', 'detection rate', 'true positive',
|
||
];
|
||
|
||
// Short/ambiguous metric tokens (ADR-263 F11): 'map' is usually the English
|
||
// word or a file extension, 'f1'/'o1' collide with finding/option labels.
|
||
// They only count as metric mentions when word-bounded, not a `.map` file
|
||
// reference, and the line (after scrubbing) carries a number — "mAP 62.3" is
|
||
// a claim, "F-numbers map to findings" is not.
|
||
// 'map' additionally must not be a `.map` file suffix or a hyphenated
|
||
// compound ("map-free", "map-reduce") — mAP the metric never appears as either.
|
||
const METRIC_TERMS_SHORT = [/(?<![.\w])map\b(?!-)/, /\bf1\b/, /\bauc\b/, /\biou\b/];
|
||
// Finding/option labels (F1, O2, …) count as labels unless the token sits in a
|
||
// metric context: an immediately following score/=/%/digit or colon ("F1: 0.91"),
|
||
// or a number later in the same clause ("F1 reaches 0.91" — an F1-score claim).
|
||
// Bare option refs ("F7 fixes", "O1–O9", "ADR-263 O2") carry no clause number of
|
||
// their own and stay labels. (A surviving 'f1' still only fires as a metric when
|
||
// its scrubbed line actually carries a number — see mentionsMetricTerm.)
|
||
const LABEL_TOKEN_RE = /\b[fo]\d+\b(?!\s*(?:score|=|\d|%|:))(?![^\n.;]*\d)/g;
|
||
const CODE_SPAN_RE = /`[^`]*`/g; // backticked identifiers are code, not claims
|
||
const HAS_NUMBER_RE = /\d/;
|
||
|
||
/** Line with code spans and finding/option labels removed. */
|
||
function scrubLine(lower) {
|
||
return lower.replace(CODE_SPAN_RE, ' ').replace(LABEL_TOKEN_RE, ' ');
|
||
}
|
||
|
||
function mentionsMetricTerm(lower, scrubbed) {
|
||
if (METRIC_TERMS.some((t) => lower.includes(t))) return true;
|
||
if (!HAS_NUMBER_RE.test(scrubbed)) return false;
|
||
return METRIC_TERMS_SHORT.some((re) => re.test(scrubbed));
|
||
}
|
||
|
||
/** Tags that make a claim honest (case-insensitive). */
|
||
const HONEST_TAGS = ['measured', 'claimed', 'synthetic', 'unvalidated', 'baseline'];
|
||
|
||
/** Reproducer references that count as evidence backing a MEASURED claim. */
|
||
const REPRODUCER_HINTS = [
|
||
'verify.py', 'witness', 'mean-pose', 'mean pose', 'held-out', 'held out',
|
||
'baseline', 'reproduce', 'sha256', 'boot log', 'pck@20 vs', 'expected_features',
|
||
// Packaging-claim reproducers (ADR-263/264 npm reviews): the tarball itself.
|
||
'npm pack', 'npm view', 'npm i ', 'npm install', 'tarball', 'cargo test',
|
||
];
|
||
|
||
const PERCENT_RE = /\b(\d{1,3}(?:\.\d+)?)\s?%/g;
|
||
// "perfect" / "100%" framing is the specific retracted claim — always high severity.
|
||
// NOTE: no trailing \b after "%": "%"→" " is non-word→non-word, so a trailing \b
|
||
// never matches and would silently miss "100%". Bare 100% is only damning next to a
|
||
// metric term (see claimCheck); the word phrases are inherently accuracy claims.
|
||
const PERFECT_PCT_RE = /\b100(?:\.0+)?\s?%/;
|
||
const PERFECT_WORD_RE = /perfect accuracy|flawless|never (?:wrong|fails)/i;
|
||
|
||
/**
|
||
* Lint a block of text for untagged or overstated accuracy claims.
|
||
* @param {string} text
|
||
* @returns {{ok: boolean, findings: Array<{severity:'high'|'medium', line:number, excerpt:string, reason:string, suggestion:string}>}}
|
||
*/
|
||
export function claimCheck(text) {
|
||
const findings = [];
|
||
if (typeof text !== 'string' || text.length === 0) {
|
||
return { ok: true, findings };
|
||
}
|
||
const lines = text.split(/\r?\n/);
|
||
|
||
lines.forEach((raw, i) => {
|
||
const line = raw.trim();
|
||
if (!line) return;
|
||
const lower = line.toLowerCase();
|
||
|
||
const hasPercent = PERCENT_RE.test(line);
|
||
PERCENT_RE.lastIndex = 0; // reset stateful global regex
|
||
const scrubbed = scrubLine(lower);
|
||
const mentionsMetric = mentionsMetricTerm(lower, scrubbed);
|
||
if (!hasPercent && !mentionsMetric) return;
|
||
|
||
const tagged = HONEST_TAGS.some((t) => lower.includes(t));
|
||
const hasReproducer = REPRODUCER_HINTS.some((h) => lower.includes(h));
|
||
const perfect = PERFECT_WORD_RE.test(line) || (mentionsMetric && PERFECT_PCT_RE.test(line));
|
||
|
||
if (perfect && !lower.includes('retract')) {
|
||
findings.push({
|
||
severity: 'high',
|
||
line: i + 1,
|
||
excerpt: clip(line),
|
||
reason: 'States perfect/100% accuracy — this is the exact framing the project retracted.',
|
||
suggestion: 'Replace with a held-out number vs the mean-pose baseline, tagged MEASURED, or mark the old claim "retracted".',
|
||
});
|
||
return;
|
||
}
|
||
|
||
// A quantitative claim needs a number. Digits hidden in a code span still
|
||
// count — "accuracy reached `0.95`" is a claim — so test the line with only
|
||
// finding/option labels stripped, NOT the code-span-scrubbed copy: scrubbing
|
||
// dropped `0.95` and wrongly short-circuited both the untagged and the
|
||
// MEASURED-without-reproducer checks below. A bare metric word in prose
|
||
// ("precision matters here", "every accuracy number must be MEASURED") has no
|
||
// number and is not a taggable claim (ADR-263 F11).
|
||
if (!hasPercent && !HAS_NUMBER_RE.test(lower.replace(LABEL_TOKEN_RE, ' '))) return;
|
||
|
||
// A metric/percent with no honesty tag at all.
|
||
if (!tagged) {
|
||
findings.push({
|
||
severity: 'medium',
|
||
line: i + 1,
|
||
excerpt: clip(line),
|
||
reason: 'Accuracy claim is not tagged MEASURED / CLAIMED / SYNTHETIC.',
|
||
suggestion: 'Tag it. If MEASURED, name the reproducer (verify.py, witness bundle, held-out vs mean-pose).',
|
||
});
|
||
return;
|
||
}
|
||
|
||
// Tagged MEASURED but cites no reproducer — still a gap (reached now even
|
||
// when the only number is inside a code span, e.g. "accuracy `0.97` (MEASURED)").
|
||
if (lower.includes('measured') && !hasReproducer) {
|
||
findings.push({
|
||
severity: 'medium',
|
||
line: i + 1,
|
||
excerpt: clip(line),
|
||
reason: 'Tagged MEASURED but cites no reproducer/evidence.',
|
||
suggestion: 'Add the evidence path: verify.py VERDICT, witness bundle, or held-out PCK vs the mean-pose baseline.',
|
||
});
|
||
}
|
||
});
|
||
|
||
return { ok: findings.length === 0, findings };
|
||
}
|
||
|
||
function clip(s, n = 120) {
|
||
return s.length > n ? `${s.slice(0, n - 1)}…` : s;
|
||
}
|
||
|
||
/** Convenience: a one-line human summary for CLI output. */
|
||
export function summarize(result) {
|
||
if (result.ok) return 'claim-check: PASS — no untagged or overstated accuracy claims.';
|
||
const high = result.findings.filter((f) => f.severity === 'high').length;
|
||
return `claim-check: ${result.findings.length} finding(s) (${high} high) — accuracy claims need MEASURED/CLAIMED tags + a reproducer.`;
|
||
}
|