mirror of
https://github.com/ruvnet/RuView
synced 2026-07-25 17:51:48 +00:00
e6f26e9ac9
* docs(adr): deep review of the RuView npm surface — ADR-263/264/265 optimization strategies
ADR-263 — @ruvnet/ruview@0.1.0 harness review (O1–O9):
- HIGH: claim-check CLI fails open on empty input (no --text/--file -> PASS exit 0)
- HIGH: MCP stdio server head-of-line blocking (spawnSync verify/calibrate up to 600s)
- MEASURED: optionalDependencies triple the cold npx install (4 pkgs/620kB/71 files
vs 1 pkg/172kB/22 files with --omit=optional) for a path that never imports them
- maxBuffer truncation, python -c port interpolation, version drift, duplicate skills,
guardrail METRIC_TERMS substring false positives ('map'/'F1' — found by dogfooding
claim-check on these very ADRs), zero CI
ADR-264 — @ruvnet/rvagent@0.1.0 + @ruv/ruview-cli review (O1–O9), verified against
the published registry tarball:
- HIGH: exports.require -> dist/index.cjs which is never built nor published
- MEASURED: 44 dead source-map files = 62,698B of the 188kB unpacked payload
- stdio-only server described as dual-transport; mixed dot/underscore tool names;
double Zod validation + hand-duplicated advertised schemas; 2-fd leak per training
job; unbounded body in the unwired HTTP scaffold; dead detectCogBinary candidates;
ruview bin-name collision
ADR-265 — cross-cutting npm distribution strategy: npm-packages.yml CI matrix
(test + pack-content/size gate + tarball-install smoke test), publish-from-CI-only
with npm provenance, version single-sourcing from package.json, bin/namespace
ownership (ruview bin belongs to @ruvnet/ruview), claim-check on package READMEs.
Docs only — no runtime code changed. Index/CHANGELOG/CLAUDE.md/README counts updated.
Co-Authored-By: claude-flow <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01WrGfTGKv1oWZ6iwXZACULz
* fix(npm): implement ADR-263/264/265 — harness fail-closed + async MCP, rvagent packaging/transport/naming, npm CI+provenance gate
ADR-263 (@ruvnet/ruview 0.2.0), O1-O9:
- claim-check fails closed on empty input (CLI exit 2, empty_text tool error)
- MCP stdio server dispatches tools/call asynchronously (promise-based spawn);
ping answers while a 3s fake verify runs — pinned by new e2e test
- optionalDependencies dropped: cold npx installs exactly 1 package
(MEASURED: was 4 pkgs/620kB/71 files via npm i in a clean prefix)
- bounded rolling output tails replace spawnSync 1MiB maxBuffer
- node_monitor port passed via sys.argv, never spliced into python -c source
- serverInfo.version read from package.json; resources/prompts stubs
- skills single-sourced: prepack sync script generates .claude/skills/ copies
- which() = memoized dep-free PATH scan
- tools underscore-canonical (ruview_claim_check, ...) + dotted aliases
- guardrail precision: word-boundary map/f1/auc/iou, code-span + F1/O2 label
scrubbing, quantitative-claims-only; packaging reproducer hints
- 30/30 tests (was 17), incl. concurrency e2e + fail-open regression pins
ADR-264 (@ruvnet/rvagent 0.2.0), O1-O9:
- exports fixed: types-first, phantom dist/index.cjs require target removed
- tarball map-free: 127,704B unpacked / 46 files / 0 maps (MEASURED,
npm pack --dry-run; was 188kB incl. 44 maps referencing unshipped src)
- Streamable HTTP actually wired behind RVAGENT_HTTP_PORT: one transport +
one MCP server per session (mcp-session-id routing), 1MiB body cap (413),
port-aware localhost origin gate; dual-transport description now true
- tools renamed underscore-canonical with dotted router-only aliases
- single Zod validation gate; advertised inputSchema generated from the same
Zod source (zod-to-json-schema)
- train_count: parent log fds closed (was leaking 2/job); job records
persisted to <jobsDir>/<id>.json (job_status survives restarts); bounded
log-tail reads
- detectCogBinary probes its candidates instead of dead-coding them
- version from package.json; @types/express dropped; @types/jest -> 29
- README rewritten to match reality (no phantom subcommands/policy layer)
- 99/99 jest tests (incl. new session/body-cap suite + previously-broken
manifest suite); stdio handshake + HTTP session flow smoke-tested live
ADR-265 D1-D4:
- .github/workflows/npm-packages.yml: 3-package x Node 20/22 gate — tests,
version-literal grep (D3), pack-content/size gate, tarball-install smoke
test (catches the ADR-264 F1 class), README claim-check (D4)
- .github/workflows/ruview-npm-release.yml: publish from CI only with
npm publish --provenance
- @ruv/ruview-cli bin renamed ruview-cli (ruview bin belongs to
@ruvnet/ruview); version single-sourced
- ci.yml NODE_VERSION 18 -> 20
ADR statuses updated to Accepted/implemented; harness manifest re-pinned;
ADR-263/264/265 + both package READMEs pass claim-check.
Co-Authored-By: claude-flow <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01WrGfTGKv1oWZ6iwXZACULz
* perf(rvagent): lazy-load HTTP transport + memoize generated tool schemas
stdio time-to-first-response ~242ms -> ~189ms (-22%; MEASURED, median of
repeated initialize round-trips against dist/index.js in this container).
- ./http-transport.js now imported lazily inside the RVAGENT_HTTP_PORT
branch: it chain-loads the MCP SDK streamableHttp module (~48ms MEASURED
via per-module import() timing) which the default stdio path never uses
- toolInputJsonSchema memoized per tool: schemas are static for the process
lifetime; under the session-per-server HTTP model every session calls
tools/list, so stop re-walking the Zod tree each time
No behavior change: 99/99 jest tests; HTTP session flow re-smoke-tested
through the lazy import path (initialize -> 200 + mcp-session-id).
Profiled @ruvnet/ruview too and left it alone: 50ms CLI startup vs ~29ms
bare 'node -e ""' floor on the same box (MEASURED) — already near the
interpreter floor with zero dependencies.
Co-Authored-By: claude-flow <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01WrGfTGKv1oWZ6iwXZACULz
* ci(ruview-cli): pass jest --passWithNoTests so the private no-test package doesn't fail the npm-packages matrix
Co-Authored-By: claude-flow <ruv@ruv.net>
* fix(npm): address 10 verified review findings in harness + rvagent before 0.2.0 publish
harness/ruview (@ruvnet/ruview):
- guardrails: digit gate now sees numbers inside code spans; F1-style
metric tokens followed by ':' or a nearby number are no longer scrubbed
(fail-open regressions in the honesty gate)
- mcp-server: tools/call requests serialize through a FIFO promise chain
(hardware/mutating tools never overlap) while ping/tools/list stay
immediate; stdin close drains in-flight responses before exit
- tools: which() no longer memoizes negative lookups
tools/ruview-mcp (@ruvnet/rvagent):
- index: realpath invoked-directly guard — library import no longer
connects a stdio transport to the consumer's process
- http-transport: explicit allowedOrigins is exact-match only (localhost
any-port convenience applies only with no configured allowlist);
session map gains maxSessions=64 + 5min idle TTL sweep
- train-count: job records persist the child pid and reconcile stale
'running' status after a server restart (exit-code marker or dead pid)
- config: cog binary candidates ordered by process.arch
.github/workflows/ruview-npm-release.yml: port the full ADR-265 D1 gate
(version-literal check, unpacked-size budget, tarball-install smoke test)
from npm-packages.yml so the publish path enforces what the header claims.
Tests: harness 30→36, rvagent 99→112, all passing.
Co-Authored-By: claude-flow <ruv@ruv.net>
---------
Co-authored-by: Claude <noreply@anthropic.com>
342 lines
11 KiB
TypeScript
342 lines
11 KiB
TypeScript
/**
|
|
* MCP tool: ruview_train_count + ruview_job_status
|
|
*
|
|
* Kick off a cog-person-count training run and poll its status.
|
|
*
|
|
* The training pipeline used here is the Candle GPU trainer from
|
|
* `v2/crates/wifi-densepose-train` — the same one that produced
|
|
* `count_v1.safetensors` in 2.1 s on the RTX 5080 (ADR-103).
|
|
*
|
|
* The MCP server shells out to `cargo run -p wifi-densepose-train --` with the
|
|
* paired JSONL path as input, redirecting stdout/stderr to a log file. The
|
|
* returned job_id can be used with ruview_job_status to poll progress.
|
|
*
|
|
* M1: job is enqueued (background process spawned, log file created).
|
|
* M4: full training arguments + real output artifact path returned.
|
|
*/
|
|
|
|
import { z } from "zod";
|
|
import { randomUUID } from "node:crypto";
|
|
import {
|
|
mkdirSync,
|
|
appendFileSync,
|
|
openSync,
|
|
closeSync,
|
|
readFileSync,
|
|
writeFileSync,
|
|
statSync,
|
|
readSync,
|
|
} from "node:fs";
|
|
import path from "node:path";
|
|
import { spawn } from "node:child_process";
|
|
import type { RuviewConfig, TrainJobResult, JobStatusResult } from "../types.js";
|
|
|
|
export const trainCountSchema = z.object({
|
|
/**
|
|
* Path to the paired JSONL file for training.
|
|
* Produced by scripts/align-ground-truth.js.
|
|
* E.g. data/paired/wiflow-p7-2026-05-19.paired.jsonl
|
|
*/
|
|
paired_jsonl: z
|
|
.string()
|
|
.describe("Absolute or relative path to the paired JSONL training file."),
|
|
/** Number of training epochs (default: 400, matching ADR-103 recipe). */
|
|
epochs: z
|
|
.number()
|
|
.int()
|
|
.min(1)
|
|
.max(10_000)
|
|
.optional()
|
|
.default(400)
|
|
.describe("Training epochs (default: 400)."),
|
|
/**
|
|
* Learning rate. The ADR-103 recipe uses 1e-3 with frozen encoder for the
|
|
* first 50 epochs, then 1e-4 for joint fine-tuning.
|
|
*/
|
|
learning_rate: z
|
|
.number()
|
|
.optional()
|
|
.default(1e-3)
|
|
.describe("Initial learning rate (default: 0.001)."),
|
|
/** Directory where the trained model artifacts are written. */
|
|
output_dir: z
|
|
.string()
|
|
.optional()
|
|
.describe(
|
|
"Directory for model artifacts (default: v2/crates/cog-person-count/cog/artifacts/)."
|
|
),
|
|
});
|
|
|
|
export type TrainCountInput = z.infer<typeof trainCountSchema>;
|
|
|
|
export const jobStatusSchema = z.object({
|
|
job_id: z.string().uuid().describe("Job ID returned by ruview_train_count."),
|
|
});
|
|
|
|
export type JobStatusInput = z.infer<typeof jobStatusSchema>;
|
|
|
|
interface JobRecord {
|
|
status: "queued" | "running" | "done" | "failed" | "unknown";
|
|
log_path: string;
|
|
queued_at: number;
|
|
epochs_total: number;
|
|
/**
|
|
* OS pid of the training child. Persisted so a later process (e.g. after an
|
|
* MCP server restart) can tell whether a job still marked 'running' actually
|
|
* outlived the process that spawned it (ADR-264 O6).
|
|
*/
|
|
pid?: number | undefined;
|
|
/** Human-readable explanation attached during reconciliation (unknown state). */
|
|
reason?: string | undefined;
|
|
}
|
|
|
|
// In-process job registry, mirrored to <jobsDir>/<id>.json on every state
|
|
// change so ruview_job_status survives an MCP server restart (ADR-264 O6).
|
|
const jobRegistry = new Map<string, JobRecord>();
|
|
|
|
function jobRecordPath(jobsDir: string, jobId: string): string {
|
|
return path.join(jobsDir, `${jobId}.json`);
|
|
}
|
|
|
|
function persistJob(jobsDir: string, jobId: string, record: JobRecord): void {
|
|
try {
|
|
writeFileSync(
|
|
jobRecordPath(jobsDir, jobId),
|
|
JSON.stringify({ job_id: jobId, ...record }, null, 2)
|
|
);
|
|
} catch {
|
|
// Persistence is best-effort; the in-memory record still serves this process.
|
|
}
|
|
}
|
|
|
|
function loadPersistedJob(jobsDir: string, jobId: string): JobRecord | undefined {
|
|
try {
|
|
const raw = JSON.parse(readFileSync(jobRecordPath(jobsDir, jobId), "utf8")) as
|
|
Partial<JobRecord>;
|
|
if (typeof raw.log_path !== "string" || typeof raw.status !== "string") {
|
|
return undefined;
|
|
}
|
|
return {
|
|
status: raw.status,
|
|
log_path: raw.log_path,
|
|
queued_at: typeof raw.queued_at === "number" ? raw.queued_at : 0,
|
|
epochs_total: typeof raw.epochs_total === "number" ? raw.epochs_total : 0,
|
|
pid: typeof raw.pid === "number" ? raw.pid : undefined,
|
|
reason: typeof raw.reason === "string" ? raw.reason : undefined,
|
|
};
|
|
} catch {
|
|
return undefined;
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Is `pid` still a live process? `process.kill(pid, 0)` sends no signal but
|
|
* probes existence: ESRCH ⇒ gone; EPERM ⇒ alive but owned by another user
|
|
* (treated as alive so we never falsely reconcile a still-running job).
|
|
*/
|
|
function isProcessAlive(pid: number): boolean {
|
|
try {
|
|
process.kill(pid, 0);
|
|
return true;
|
|
} catch (e) {
|
|
return (e as NodeJS.ErrnoException).code === "EPERM";
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Scan log lines (tail) for the "# exit code: N" marker the child.on('close')
|
|
* handler appends. `found:false` means the process died without the marker —
|
|
* i.e. this server never saw the close (it restarted mid-run).
|
|
*/
|
|
function findExitMarker(lines: string[]): { found: boolean; code: number | null } {
|
|
for (let i = lines.length - 1; i >= 0; i--) {
|
|
const m = /^# exit code: (-?\d+|null)$/.exec((lines[i] ?? "").trim());
|
|
if (m) return { found: true, code: m[1] === "null" ? null : Number(m[1]) };
|
|
}
|
|
return { found: false, code: null };
|
|
}
|
|
|
|
/** Read the last `maxLines` lines of a file without loading the whole log. */
|
|
function tailLines(filePath: string, maxLines: number, maxBytes = 64 * 1024): string[] {
|
|
const size = statSync(filePath).size;
|
|
const start = Math.max(0, size - maxBytes);
|
|
const buf = Buffer.alloc(size - start);
|
|
const fd = openSync(filePath, "r");
|
|
try {
|
|
readSync(fd, buf, 0, buf.length, start);
|
|
} finally {
|
|
closeSync(fd);
|
|
}
|
|
const lines = buf.toString("utf8").split("\n");
|
|
return lines.slice(Math.max(0, lines.length - maxLines));
|
|
}
|
|
|
|
export async function trainCount(
|
|
input: TrainCountInput,
|
|
config: RuviewConfig
|
|
): Promise<object> {
|
|
const jobId = randomUUID();
|
|
const logDir = config.jobsDir;
|
|
mkdirSync(logDir, { recursive: true });
|
|
const logPath = path.join(logDir, `${jobId}.log`);
|
|
const queuedAt = Date.now() / 1000;
|
|
|
|
// Default output directory matches ADR-103 repo layout.
|
|
const outputDir =
|
|
input.output_dir ?? "v2/crates/cog-person-count/cog/artifacts";
|
|
|
|
// Record the job immediately so ruview_job_status can find it — in memory
|
|
// and on disk (survives server restarts, ADR-264 O6).
|
|
const record: JobRecord = {
|
|
status: "queued",
|
|
log_path: logPath,
|
|
queued_at: queuedAt,
|
|
epochs_total: input.epochs,
|
|
};
|
|
jobRegistry.set(jobId, record);
|
|
persistJob(logDir, jobId, record);
|
|
|
|
// Write the header synchronously so the log file exists before spawn.
|
|
const header = [
|
|
`# RuView training job ${jobId}`,
|
|
`# started: ${new Date().toISOString()}`,
|
|
`# paired_jsonl: ${input.paired_jsonl}`,
|
|
`# epochs: ${input.epochs}`,
|
|
`# learning_rate: ${input.learning_rate}`,
|
|
`# output_dir: ${outputDir}`,
|
|
"",
|
|
].join("\n");
|
|
appendFileSync(logPath, header);
|
|
|
|
// Open log file descriptors synchronously (avoids WriteStream-before-open bug on Windows).
|
|
const logFdOut = openSync(logPath, "a");
|
|
const logFdErr = openSync(logPath, "a");
|
|
|
|
const args = [
|
|
"run",
|
|
"--release",
|
|
"-p",
|
|
"wifi-densepose-train",
|
|
"--",
|
|
"--task",
|
|
"count",
|
|
"--paired",
|
|
input.paired_jsonl,
|
|
"--epochs",
|
|
String(input.epochs),
|
|
"--lr",
|
|
String(input.learning_rate),
|
|
"--output-dir",
|
|
outputDir,
|
|
];
|
|
|
|
// M1: cargo may not be on PATH on non-Rust machines — spawn fails gracefully.
|
|
const child = spawn("cargo", args, {
|
|
detached: true,
|
|
stdio: ["ignore", logFdOut, logFdErr],
|
|
});
|
|
|
|
child.unref(); // Allow the MCP server process to exit without waiting for training.
|
|
|
|
// The child holds its own duplicates of the log fds; close the parent's
|
|
// copies immediately or every job leaks 2 fds for the server's lifetime
|
|
// (ADR-264 F6/O6).
|
|
closeSync(logFdOut);
|
|
closeSync(logFdErr);
|
|
|
|
// Record the child pid so a later process can reconcile a stale 'running'
|
|
// record after a server restart (child.pid is undefined only if spawn failed
|
|
// synchronously, in which case the 'error' handler flips status to 'failed').
|
|
record.pid = child.pid;
|
|
record.status = "running";
|
|
persistJob(logDir, jobId, record);
|
|
|
|
child.on("error", (e) => {
|
|
appendFileSync(logPath, `\n# ERROR: ${e.message}\n`);
|
|
record.status = "failed";
|
|
persistJob(logDir, jobId, record);
|
|
});
|
|
|
|
child.on("close", (code) => {
|
|
appendFileSync(logPath, `\n# exit code: ${code}\n`);
|
|
record.status = code === 0 ? "done" : "failed";
|
|
persistJob(logDir, jobId, record);
|
|
});
|
|
|
|
const result: TrainJobResult = {
|
|
job_id: jobId,
|
|
status: "running",
|
|
log_path: logPath,
|
|
queued_at: queuedAt,
|
|
};
|
|
|
|
return {
|
|
ok: true,
|
|
result,
|
|
note:
|
|
"Training job spawned in the background. " +
|
|
`Poll progress with ruview_job_status({ job_id: "${jobId}" }). ` +
|
|
`Live log: ${logPath}`,
|
|
};
|
|
}
|
|
|
|
export async function jobStatus(
|
|
input: JobStatusInput,
|
|
config: RuviewConfig
|
|
): Promise<object> {
|
|
// Memory first, then the persisted record (survives server restarts).
|
|
let job = jobRegistry.get(input.job_id) ?? loadPersistedJob(config.jobsDir, input.job_id);
|
|
if (!job) {
|
|
return {
|
|
ok: false,
|
|
error: `Job ${input.job_id} not found in this server or in ${config.jobsDir}.`,
|
|
};
|
|
}
|
|
|
|
// Reconcile a 'running' record whose owning process is gone. The status flip
|
|
// to done/failed lives only in the spawning process's child.on('close'/'error')
|
|
// handlers; if this server restarted mid-run, the record froze at 'running'
|
|
// (ADR-264 O6). When the pid is dead, recover the true outcome from the log's
|
|
// "# exit code: N" marker, else surface an honest 'unknown'.
|
|
if (job.status === "running" && typeof job.pid === "number" && !isProcessAlive(job.pid)) {
|
|
let tail: string[] = [];
|
|
try {
|
|
tail = tailLines(job.log_path, 40);
|
|
} catch {
|
|
/* log unreadable — treated as no marker below */
|
|
}
|
|
const marker = findExitMarker(tail);
|
|
const reconciled: JobRecord = { ...job };
|
|
if (marker.found) {
|
|
reconciled.status = marker.code === 0 ? "done" : "failed";
|
|
reconciled.reason = undefined;
|
|
} else {
|
|
reconciled.status = "unknown";
|
|
reconciled.reason =
|
|
"process gone, no exit marker — server likely restarted mid-run";
|
|
}
|
|
jobRegistry.set(input.job_id, reconciled);
|
|
persistJob(config.jobsDir, input.job_id, reconciled);
|
|
job = reconciled;
|
|
}
|
|
|
|
// Bounded tail read — never load a multi-GB training log wholesale.
|
|
let recentLog: string[] = [];
|
|
try {
|
|
recentLog = tailLines(job.log_path, 20);
|
|
} catch {
|
|
recentLog = ["(log not readable yet)"];
|
|
}
|
|
|
|
const result: JobStatusResult = {
|
|
job_id: input.job_id,
|
|
status: job.status,
|
|
log_path: job.log_path,
|
|
recent_log: recentLog,
|
|
epochs_total: job.epochs_total,
|
|
...(job.reason !== undefined ? { reason: job.reason } : {}),
|
|
};
|
|
|
|
return { ok: true, result };
|
|
}
|