diff --git a/benchmarks/wiflow-std/.gitignore b/benchmarks/wiflow-std/.gitignore index 6d608ac1..56133dbc 100644 --- a/benchmarks/wiflow-std/.gitignore +++ b/benchmarks/wiflow-std/.gitignore @@ -16,3 +16,4 @@ downloads/ *.safetensors results/parity_fixture.json __pycache__/ +*.onnx diff --git a/benchmarks/wiflow-std/RESULTS.md b/benchmarks/wiflow-std/RESULTS.md index 2c1cda6c..48f38933 100644 --- a/benchmarks/wiflow-std/RESULTS.md +++ b/benchmarks/wiflow-std/RESULTS.md @@ -104,7 +104,113 @@ checkpoint REFUTED; dataset/code require repairs)**. RuView docs may cite ESP32 numbers (different hardware, 5 subjects, in-domain random split, 15 keypoints). +## Edge optimization (measured) + +ADR-152 "optimize beyond SOTA" track, 2026-06-10, this Windows box (Windows 11, +16 torch threads, torch 2.12.0+cpu, onnxruntime 1.26.0). Subject: the retrained +checkpoint `results/retrained_best_pose_model.pth` (2,225,042 fp32 params). +Scripts: `quantize_bench.py`, `onnx_bench.py`, `eval_ort_accuracy.py`. +Raw numbers: `results/edge_optimization.json`. + +Accuracy is on a **10,000-window seed-42 random subset** of the corruption-free +test split (same seed-42 file-level 70/15/15 split as `eval_repro.py`; 54,000 +test windows, 1,440 corrupted excluded via `results/nan_windows_mask.npy` | +`results/big_windows_mask.npy`, leaving 52,560; subset drawn with +`np.random.default_rng(42)`). The fp32 subset PCK@20 (96.68%) matches the full +clean-test figure (96.61%), so the subset is representative. + +Latency is CPU ms/window, median of repeated runs, 3 interleaved repetitions +per variant (medians below; run-to-run spread on this box is large, roughly +±20-40% at batch 1 — reps are in the JSON). + +| Variant | Disk size | Batch 1 (ms/win) | Batch 64 (ms/win) | PCK@20 | PCK@50 | MPJPE | +|---|---|---|---|---|---|---| +| torch fp32 (baseline) | 9.07 MB | 11.0 | 2.27 | 96.68% | 99.15% | 0.00936 | +| torch fp16 (`.half()`) | **4.58 MB** | 24.3 | 2.42 | 96.68% | 99.15% | 0.00946 | +| torch int8 dynamic | 9.07 MB (unchanged) | 15.6 | 2.06 | 96.68% (identical) | 99.15% | 0.00936 | +| ONNX fp32 (onnxruntime) | 8.97 MB | **3.2** | **2.0** | 96.68% | 99.15% | 0.00936 | +| ONNX int8 (ORT dynamic, supplementary) | **2.44 MB** | 6.5 | 5.8 | 96.52% | 99.15% | 0.01108 | + +Findings: + +- **torch dynamic INT8 quantizes nothing on this model.** The architecture has + **zero `nn.Linear` layers** — it is entirely Conv1d (21) + Conv2d (22) + + BatchNorm. `torch.ao.quantization.quantize_dynamic` (requested over + `{Linear, Conv1d, Conv2d}`) converted **0 modules / 0.0% of params**: dynamic + quantization only has kernels for Linear/RNN-family modules and silently + skips convolutions. The "int8" model is bit-identical to fp32 (same outputs, + same 9.07 MB). Conv quantization would require static (PTQ) quantization + with calibration — out of scope here; the ORT dynamic path below is the + honest int8 datapoint. +- **fp16 halves size for free accuracy-wise** (PCK@20 −0.005 pt, MPJPE + +0.0001) but is *slower* on CPU at batch 1 (~2.2×) — torch CPU fp16 conv + kernels are emulated. fp16 is a storage/transport format here, not a CPU + runtime win. +- **ONNX Runtime is the real batch-1 latency win: ~3.4× faster than torch** + (3.2 vs 11.0 ms/window) at identical accuracy (parity 2.4e-7). + +### Verdict on the paper's "~2.2 MB int8" claim + +**Plausible but not free, and unreachable by the obvious PyTorch route.** +2,225,042 params × 1 byte ≈ 2.2 MB assumes *every* parameter quantizes. +PyTorch dynamic quantization — the one-liner most readers would reach for — +yields **9.07 MB (0% quantized)** because the model has no Linear layers. +ONNX Runtime dynamic quantization, which does have int8 conv weight support, +gets **2.44 MB** (close to the claim; the overhead is BatchNorm params/buffers +and quantization scales kept in fp32) at a measurable accuracy cost: +PCK@20 96.68 → 96.52% (−0.16 pt) and MPJPE 0.00936 → 0.01108 (+18%), and +~2× slower inference than ONNX fp32 (ConvInteger kernels). The paper does not +state a method or an int8 accuracy; treat "2.2 MB" as a weight-arithmetic +estimate, achievable in practice only via conv-capable quantization toolchains +and with a small accuracy penalty. + +### ONNX export status + +**Works.** Exported via the TorchScript exporter (`dynamo=False`), opset 17, +with a dynamic batch axis — `results/retrained_fp32_dynamic.onnx` (8.97 MB), +verified to run at batch 1/2/64. The axial attention's +`view(N*W, C, H)` reshape traced correctly (sizes recorded as graph ops, not +baked constants). The dynamo exporter also captures the graph but crashed on +this box writing a ✅ to a cp1252 console (cosmetic Windows encoding issue, not +a model blocker). Parity vs torch on the stored fixture +(`results/parity_fixture.npz`, batch 2, seed 42): **max abs diff 2.4e-7 — +PASS** (< 1e-4). ORT-quantized int8 model: `results/retrained_int8_ort_dynamic.onnx`. + +## Measurement (b): BLOCKED-ON-DATA (attempted 2026-06-10) + +The fine-tune-on-ESP32 measurement stopped at dataset characterization, per the +pre-registered stop rule (<2,000 paired windows). Findings (MEASURED): + +- **Only one trainable paired dataset exists**: `ruvultra:~/work/cog-pose-train/paired.jsonl` + — 1,077 windows (one subject, one room, one 29.9-min session, single node; + CSI [56, 20]; 17 COCO keypoints, MediaPipe confidence mean 0.44 — only 264 + windows pass ADR-079's own conf>0.5 training filter). Prior measured attempts + on this exact set: 0–3% torso-PCK@20 (temporal splits, three independent + pipelines). Fine-tuning a 2.23M-param model on ~860 train windows would + measure memorization, not transfer. +- **The April session behind the old "92.9% PCK@20" claim is lost** (345 + samples, 35 subcarriers; raw CSI gone from ruvzen/ruvultra/cognitum-v0; only + a 69-sample predictions+GT holdout survives at `models/wiflow-real/eval-holdout.jsonl`). +- **Forensic recheck of that holdout RETRACTS the 92.9% figure**: the trainer's + `pck()` used an absolute 0.2 image-unit threshold (not torso-normalized) and + the model output a **constant pose** (pred std 0.0000 across 69 near-static + frames; a mean predictor scores 100% under the same protocol). The + torso-normalized PCK@20 on the same holdout is 19.1%. This corroborates the + 2026-05-11 audit retraction (CHANGELOG, PR #535); stale doc citations were + removed 2026-06-10 (user-guide, readme-details, ADR-152 §2.1.3). The §2.2 + no-citation rule now applies to ADR-079 accuracy claims. + +Unblock criteria: a paired collection session of ≥2k windows (≈35+ min at the +observed stride; multi-pose, conf>0.5, ideally with the §2.1.3 two-checkerboard +calibration), plus a re-baselined our-pipeline number under torso-PCK@20 on the +same split. WiFlow-STD assets stand ready on ruvultra (`~/wiflow-std-bench/`). +Also worth investigating: ADR-079's protocol predicts ~9k windows per 30 min; +the May session under-delivered ~8× (aligner drop rate?). + ## Pending -- (b) fine-tune on our ESP32 17-keypoint eval set. -- (c) our internal WiFlow on their dataset (15-keypoint subset mapping). +- (b) fine-tune on our ESP32 17-keypoint eval set — **BLOCKED-ON-DATA**, see above. +- (c) our internal WiFlow on their dataset (15-keypoint subset mapping) — also + affected: there is currently no validated internal pose model to compare + (the 92.9% artifact is retracted; the MM-Fi SOTA models in ADR-150 §3 are a + different input domain). diff --git a/benchmarks/wiflow-std/eval_ort_accuracy.py b/benchmarks/wiflow-std/eval_ort_accuracy.py new file mode 100644 index 00000000..280e627f --- /dev/null +++ b/benchmarks/wiflow-std/eval_ort_accuracy.py @@ -0,0 +1,91 @@ +"""ADR-152 edge optimization: accuracy of the ONNX fp32 and ORT-dynamic-int8 +models on the same corruption-free 10k test subset used by quantize_bench.py. + +The torch dynamic-int8 path quantizes nothing (no nn.Linear in the model), so +the only real int8 datapoint for the paper's "~2.2 MB int8" claim is the +onnxruntime dynamically quantized model -- this script measures what that +quantization costs in PCK/MPJPE. + +Usage: + .venv/Scripts/python.exe eval_ort_accuracy.py \ + --data-dir [--subset 10000] + +Writes/merges into results/edge_optimization.json under key "onnx_accuracy". +""" + +import argparse +import json +import os +import sys +import time + +import numpy as np +import torch + +HERE = os.path.dirname(os.path.abspath(__file__)) +RESULTS = os.path.join(HERE, "results") +sys.path.insert(0, HERE) + +from quantize_bench import build_test_subset # noqa: E402 (sets up upstream imports) + +sys.path.insert(0, os.path.join(HERE, "upstream")) +from utils.metrics import calculate_mpjpe, calculate_pck # noqa: E402 + + +def evaluate_ort(sess, loader, label): + inp = sess.get_inputs()[0].name + totals = {0.2: 0.0, 0.5: 0.0} + total_mpe, n = 0.0, 0 + t0 = time.time() + for batch_idx, (bx, by) in enumerate(loader): + out = torch.from_numpy(sess.run(None, {inp: bx.numpy()})[0]) + pck = calculate_pck(out, by, thresholds=[0.2, 0.5]) + mpe = calculate_mpjpe(out, by) + bs = by.size(0) + total_mpe += mpe * bs + for t in totals: + totals[t] += pck[t] * bs + n += bs + if batch_idx % 50 == 0: + print(f" [{label}] batch {batch_idx}: n={n} " + f"pck20={totals[0.2]/n:.4f} mpjpe={total_mpe/n:.4f} " + f"({time.time()-t0:.0f}s)", flush=True) + return {"samples": n, "pck@20": totals[0.2] / n, "pck@50": totals[0.5] / n, + "mpjpe": total_mpe / n, "wall_seconds": time.time() - t0} + + +def main(): + import onnxruntime as ort + parser = argparse.ArgumentParser() + parser.add_argument("--data-dir", default=os.path.join( + os.path.expanduser("~"), ".cache", "kagglehub", "datasets", "kaka2434", + "wiflow-dataset", "versions", "1", "preprocessed_csi_data")) + parser.add_argument("--subset", type=int, default=10000) + parser.add_argument("--out", default=os.path.join(RESULTS, "edge_optimization.json")) + args = parser.parse_args() + + loader, _n_clean = build_test_subset(args.data_dir, args.subset) + results = {} + for label, fname in (("onnx_fp32", "retrained_fp32_dynamic.onnx"), + ("onnx_int8_ort_dynamic", "retrained_int8_ort_dynamic.onnx")): + path = os.path.join(RESULTS, fname) + if not os.path.exists(path): + results[label] = {"error": f"{fname} not found; run onnx_bench.py first"} + continue + sess = ort.InferenceSession(path, providers=["CPUExecutionProvider"]) + print(f"=== accuracy: {label} ({fname}) ===") + results[label] = evaluate_ort(sess, loader, label) + print(json.dumps(results[label], indent=2)) + + merged = {} + if os.path.exists(args.out): + with open(args.out) as f: + merged = json.load(f) + merged["onnx_accuracy"] = results + with open(args.out, "w") as f: + json.dump(merged, f, indent=2) + print(f"wrote {args.out}") + + +if __name__ == "__main__": + main() diff --git a/benchmarks/wiflow-std/onnx_bench.py b/benchmarks/wiflow-std/onnx_bench.py new file mode 100644 index 00000000..5e4c7c14 --- /dev/null +++ b/benchmarks/wiflow-std/onnx_bench.py @@ -0,0 +1,228 @@ +"""ADR-152 edge optimization: ONNX export + onnxruntime CPU benchmark for the +retrained WiFlow-STD checkpoint. + +- Exports fp32 to ONNX. The axial attention reshapes with python ints taken + from tensor.size() (view(N*W, C, H)), so a traced graph bakes the batch + size; we first try a dynamic-batch export and verify it actually works at + batch sizes 1/2/64 -- if not, we fall back to fixed-batch exports. +- Verifies output parity vs torch on the stored fixture + (results/parity_fixture.npz, batch 2, seed 42): max abs diff < 1e-4. +- Measures onnxruntime CPU latency at batch 1 and 64 (median of N runs). +- Supplementary: onnxruntime dynamic int8 quantization of the exported model + (weight size datapoint for the paper's "~2.2 MB int8" claim). + +Usage: + .venv/Scripts/python.exe onnx_bench.py + +Writes/merges into results/edge_optimization.json under key "onnx". +""" + +import json +import os +import platform +import statistics +import sys +import time +import traceback + +import numpy as np +import torch + +HERE = os.path.dirname(os.path.abspath(__file__)) +UPSTREAM = os.path.join(HERE, "upstream") +RESULTS = os.path.join(HERE, "results") +sys.path.insert(0, UPSTREAM) + +import types # noqa: E402 + +_models_pkg = types.ModuleType("models") +_models_pkg.__path__ = [os.path.join(UPSTREAM, "models")] +sys.modules["models"] = _models_pkg + +from models.pose_model import WiFlowPoseModel # noqa: E402 + +CHECKPOINT = os.path.join(RESULTS, "retrained_best_pose_model.pth") +OUT_JSON = os.path.join(RESULTS, "edge_optimization.json") + + +def load_fp32_model(): + state = torch.load(CHECKPOINT, map_location="cpu", weights_only=True) + renames = {"att.": "attention.", "final_conv.": "decoder."} + state = {next((new + k[len(old):] for old, new in renames.items() + if k.startswith(old)), k): v + for k, v in state.items()} + model = WiFlowPoseModel(dropout=0.5) + model.load_state_dict(state, strict=True) + model.eval() + return model + + +def try_export(model, path, batch, dynamic, opset=17): + """Returns (ok, exporter_used, error).""" + x = torch.rand(batch, 540, 20) + attempts = [] + if dynamic: + attempts.append(("dynamo", dict(dynamo=True, + dynamic_shapes={"x": {0: "batch"}}))) + attempts.append(("torchscript", dict(dynamo=False, + dynamic_axes={"input": {0: "batch"}, + "output": {0: "batch"}}))) + else: + attempts.append(("torchscript", dict(dynamo=False))) + attempts.append(("dynamo", dict(dynamo=True))) + last_err = None + for name, kw in attempts: + try: + with torch.no_grad(): + torch.onnx.export(model, (x,), path, opset_version=opset, + input_names=["input"], output_names=["output"], + **kw) + return True, name, None + except Exception as e: # noqa: BLE001 + last_err = f"{name}: {type(e).__name__}: {e}" + traceback.print_exc() + return False, None, last_err + + +def ort_session(path): + import onnxruntime as ort + return ort.InferenceSession(path, providers=["CPUExecutionProvider"]) + + +def ort_run(sess, x): + inp = sess.get_inputs()[0].name + return sess.run(None, {inp: x})[0] + + +def bench_ort(sess, batch, n_runs): + rng = np.random.default_rng(123) + x = rng.random((batch, 540, 20), dtype=np.float32) + for _ in range(max(5, n_runs // 10)): + ort_run(sess, x) + times = [] + for _ in range(n_runs): + t0 = time.perf_counter() + ort_run(sess, x) + times.append(time.perf_counter() - t0) + med = statistics.median(times) + return { + "batch_size": batch, + "runs": n_runs, + "median_ms_per_batch": med * 1e3, + "median_ms_per_window": med * 1e3 / batch, + "windows_per_second": batch / med, + } + + +def main(): + import onnxruntime + model = load_fp32_model() + results = { + "env": { + "torch": torch.__version__, + "onnxruntime": onnxruntime.__version__, + "platform": platform.platform(), + }, + } + + fixture = np.load(os.path.join(RESULTS, "parity_fixture.npz")) + fx, fy = fixture["input"], fixture["output"] # (2,540,20) -> (2,15,2) + + # ---- export: dynamic batch first, fall back to fixed -------------------- + dyn_path = os.path.join(RESULTS, "retrained_fp32_dynamic.onnx") + ok, exporter, err = try_export(model, dyn_path, batch=2, dynamic=True) + dynamic_works = False + if ok: + # verify the dynamic graph really runs at other batch sizes + try: + sess = ort_session(dyn_path) + for b in (1, 2, 64): + y = ort_run(sess, np.zeros((b, 540, 20), dtype=np.float32)) + assert y.shape == (b, 15, 2), y.shape + dynamic_works = True + except Exception as e: # noqa: BLE001 + print(f"dynamic-batch model does not generalize: {e}") + + sessions = {} + if dynamic_works: + results["export"] = {"mode": "dynamic-batch", "exporter": exporter, + "file": os.path.basename(dyn_path), + "size_mb": os.path.getsize(dyn_path) / 1e6} + sess = ort_session(dyn_path) + sessions = {1: sess, 2: sess, 64: sess} + print(f"dynamic-batch export OK via {exporter}") + else: + results["export"] = {"mode": "fixed-batch", "fallback_reason": err, + "files": {}} + for b in (1, 2, 64): + p = os.path.join(RESULTS, f"retrained_fp32_b{b}.onnx") + ok, exporter, err = try_export(model, p, batch=b, dynamic=False) + if not ok: + results["export"]["files"][str(b)] = {"error": err} + print(f"EXPORT FAILED at batch {b}: {err}") + continue + results["export"]["files"][str(b)] = { + "exporter": exporter, "file": os.path.basename(p), + "size_mb": os.path.getsize(p) / 1e6} + sessions[b] = ort_session(p) + print(f"fixed-batch {b} export OK via {exporter}") + + # ---- parity vs torch on the fixture ------------------------------------- + if 2 in sessions: + y_ort = ort_run(sessions[2], fx) + with torch.no_grad(): + y_torch = model(torch.from_numpy(fx)).numpy() + results["parity"] = { + "fixture": "results/parity_fixture.npz (batch 2, seed 42)", + "max_abs_diff_vs_stored_fixture": float(np.abs(y_ort - fy).max()), + "max_abs_diff_vs_torch_now": float(np.abs(y_ort - y_torch).max()), + "pass_lt_1e-4": bool(np.abs(y_ort - y_torch).max() < 1e-4), + } + print("parity:", json.dumps(results["parity"], indent=2)) + + # ---- latency ------------------------------------------------------------- + results["latency"] = {} + if 1 in sessions: + results["latency"]["batch1"] = bench_ort(sessions[1], 1, 100) + print(f"ORT batch 1: {results['latency']['batch1']['median_ms_per_window']:.2f} ms/window") + if 64 in sessions: + results["latency"]["batch64"] = bench_ort(sessions[64], 64, 30) + print(f"ORT batch 64: {results['latency']['batch64']['median_ms_per_window']:.3f} ms/window") + + # ---- supplementary: ORT dynamic int8 (size datapoint for the 2.2MB claim) + src = (dyn_path if dynamic_works + else os.path.join(RESULTS, "retrained_fp32_b1.onnx")) + if os.path.exists(src): + try: + from onnxruntime.quantization import QuantType, quantize_dynamic + q_path = os.path.join(RESULTS, "retrained_int8_ort_dynamic.onnx") + quantize_dynamic(src, q_path, weight_type=QuantType.QInt8) + entry = {"file": os.path.basename(q_path), + "size_mb": os.path.getsize(q_path) / 1e6} + try: + qs = ort_session(q_path) + yq = ort_run(qs, fx[:1] if not dynamic_works else fx) + ref = fy[:1] if not dynamic_works else fy + entry["runs"] = True + entry["max_abs_diff_vs_fp32_fixture"] = float(np.abs(yq - ref).max()) + except Exception as e: # noqa: BLE001 + entry["runs"] = False + entry["run_error"] = f"{type(e).__name__}: {e}" + results["ort_int8_dynamic_supplementary"] = entry + print("ORT int8:", json.dumps(entry, indent=2)) + except Exception as e: # noqa: BLE001 + results["ort_int8_dynamic_supplementary"] = { + "error": f"{type(e).__name__}: {e}"} + + merged = {} + if os.path.exists(OUT_JSON): + with open(OUT_JSON) as f: + merged = json.load(f) + merged["onnx"] = results + with open(OUT_JSON, "w") as f: + json.dump(merged, f, indent=2) + print(f"wrote {OUT_JSON}") + + +if __name__ == "__main__": + main() diff --git a/benchmarks/wiflow-std/quantize_bench.py b/benchmarks/wiflow-std/quantize_bench.py new file mode 100644 index 00000000..70939fc6 --- /dev/null +++ b/benchmarks/wiflow-std/quantize_bench.py @@ -0,0 +1,290 @@ +"""ADR-152 "optimize beyond SOTA": edge-optimization benchmark for the +retrained WiFlow-STD checkpoint (results/retrained_best_pose_model.pth, +~96% PCK@20, fp32 params 2,225,042). + +Measures, for fp32 / fp16 / dynamic-int8 torch variants: + (a) serialized state_dict size on disk, + (b) CPU inference latency per window at batch 1 and batch 64 + (median of repeated runs, this Windows box), + (c) accuracy (PCK@20/50 + MPJPE, upstream metrics) on a corruption-free + random subset of the seed-42 file-level 70/15/15 test split + (same split as eval_repro.py; corrupted windows 487-499 excluded via + results/nan_windows_mask.npy | results/big_windows_mask.npy). + +Also verifies the paper's "~2.2 MB int8" size claim: reports which layer +types torch dynamic quantization actually converts (the model contains NO +nn.Linear -- it is Conv1d/Conv2d/BatchNorm only) and the real on-disk size. + +Usage: + .venv/Scripts/python.exe quantize_bench.py \ + --data-dir C:/Users/ruv/.cache/kagglehub/datasets/kaka2434/wiflow-dataset/versions/1/preprocessed_csi_data \ + [--subset 10000] [--skip-accuracy] + +Writes/merges into results/edge_optimization.json under key "torch". +""" + +import argparse +import json +import os +import platform +import statistics +import sys +import time + +import numpy as np +import torch +import torch.nn as nn +from torch.utils.data import DataLoader + +HERE = os.path.dirname(os.path.abspath(__file__)) +UPSTREAM = os.path.join(HERE, "upstream") +RESULTS = os.path.join(HERE, "results") +sys.path.insert(0, UPSTREAM) + +# Upstream models/__init__.py is broken as published (imports a name tcn.py +# does not define); register a stub package so it never executes. +import types # noqa: E402 + +_models_pkg = types.ModuleType("models") +_models_pkg.__path__ = [os.path.join(UPSTREAM, "models")] +sys.modules["models"] = _models_pkg + +import dataset as upstream_dataset # noqa: E402 +from dataset import ( # noqa: E402 + PreprocessedCSIKeypointsDataset, + create_preprocessed_train_val_test_loaders, +) +from models.pose_model import WiFlowPoseModel # noqa: E402 +from utils.metrics import calculate_mpjpe, calculate_pck # noqa: E402 + +CHECKPOINT = os.path.join(RESULTS, "retrained_best_pose_model.pth") + +# csi_windows.npy is ~13 GB; mmap large arrays instead of loading into RAM +# (same trick as eval_repro.py). +_np_load = np.load + + +def _np_load_mmap(path, *a, **kw): + if (isinstance(path, str) and path.endswith(".npy") + and os.path.getsize(path) > 1 << 30 and "mmap_mode" not in kw): + kw["mmap_mode"] = "r" + return _np_load(path, *a, **kw) + + +upstream_dataset.np.load = _np_load_mmap + + +def load_fp32_model(): + state = torch.load(CHECKPOINT, map_location="cpu", weights_only=True) + # legacy upstream names, harmless no-op on the retrained checkpoint + renames = {"att.": "attention.", "final_conv.": "decoder."} + state = {next((new + k[len(old):] for old, new in renames.items() + if k.startswith(old)), k): v + for k, v in state.items()} + model = WiFlowPoseModel(dropout=0.5) + model.load_state_dict(state, strict=True) + model.eval() + return model + + +def state_dict_size_bytes(model, path): + torch.save(model.state_dict(), path) + return os.path.getsize(path) + + +def bench_latency(model, batch_size, n_runs, dtype=torch.float32): + gen = torch.Generator().manual_seed(123) + x = torch.rand(batch_size, 540, 20, generator=gen).to(dtype) + with torch.no_grad(): + for _ in range(max(5, n_runs // 10)): # warmup + model(x) + times = [] + for _ in range(n_runs): + t0 = time.perf_counter() + model(x) + times.append(time.perf_counter() - t0) + med = statistics.median(times) + return { + "batch_size": batch_size, + "runs": n_runs, + "median_ms_per_batch": med * 1e3, + "median_ms_per_window": med * 1e3 / batch_size, + "windows_per_second": batch_size / med, + } + + +def build_test_subset(data_dir, subset_size, batch_size=64): + """Seed-42 file-level 70/15/15 test split (exactly as eval_repro.py), + minus corrupted windows, then a seed-42 random subset.""" + dataset = PreprocessedCSIKeypointsDataset( + data_dir=data_dir, keypoint_scale=1000.0, enable_temporal_clean=True) + _tr, _va, test_loader = create_preprocessed_train_val_test_loaders( + dataset=dataset, batch_size=batch_size, num_workers=0, random_seed=42) + test_indices = np.asarray(test_loader.dataset.indices) + + corrupted = (np.load(os.path.join(RESULTS, "nan_windows_mask.npy")) + | np.load(os.path.join(RESULTS, "big_windows_mask.npy"))) + clean = test_indices[~corrupted[test_indices]] + print(f"test split: {len(test_indices)} windows, " + f"{len(test_indices) - len(clean)} corrupted excluded, " + f"{len(clean)} clean") + + if subset_size and subset_size < len(clean): + rng = np.random.default_rng(42) + clean = np.sort(rng.choice(clean, size=subset_size, replace=False)) + subset = torch.utils.data.Subset(dataset, clean.tolist()) + loader = DataLoader(subset, batch_size=batch_size, shuffle=False, + num_workers=0) + return loader, len(clean) + + +def evaluate(model, loader, dtype=torch.float32, label=""): + totals = {0.2: 0.0, 0.5: 0.0} + total_mpe, n = 0.0, 0 + t0 = time.time() + with torch.no_grad(): + for batch_idx, (bx, by) in enumerate(loader): + out = model(bx.to(dtype)).float() + pck = calculate_pck(out, by, thresholds=[0.2, 0.5]) + mpe = calculate_mpjpe(out, by) + bs = by.size(0) + total_mpe += mpe * bs + for t in totals: + totals[t] += pck[t] * bs + n += bs + if batch_idx % 50 == 0: + print(f" [{label}] batch {batch_idx}: n={n} " + f"pck20={totals[0.2]/n:.4f} mpjpe={total_mpe/n:.4f} " + f"({time.time()-t0:.0f}s)", flush=True) + return { + "samples": n, + "pck@20": totals[0.2] / n, + "pck@50": totals[0.5] / n, + "mpjpe": total_mpe / n, + "wall_seconds": time.time() - t0, + } + + +def quantize_int8_dynamic(fp32_model): + """torch.ao.quantization.quantize_dynamic on Linear/Conv where supported. + Returns (model, report) where report documents what actually quantized.""" + qmodel = torch.ao.quantization.quantize_dynamic( + fp32_model, {nn.Linear, nn.Conv1d, nn.Conv2d}, dtype=torch.qint8) + + quantized, total_params, quant_params = [], 0, 0 + for name, mod in qmodel.named_modules(): + cls = type(mod).__module__ + "." + type(mod).__name__ + if "quantized" in cls: + w = mod.weight() if callable(getattr(mod, "weight", None)) else None + numel = w.numel() if w is not None else 0 + quant_params += numel + quantized.append({"module": name, "class": cls, "params": numel}) + for p in fp32_model.parameters(): + total_params += p.numel() + + n_linear = sum(isinstance(m, nn.Linear) for m in fp32_model.modules()) + n_conv1d = sum(isinstance(m, nn.Conv1d) for m in fp32_model.modules()) + n_conv2d = sum(isinstance(m, nn.Conv2d) for m in fp32_model.modules()) + report = { + "eligible_module_counts": { + "nn.Linear": n_linear, "nn.Conv1d": n_conv1d, "nn.Conv2d": n_conv2d}, + "modules_actually_quantized": quantized, + "n_modules_quantized": len(quantized), + "params_total": total_params, + "params_quantized": quant_params, + "params_quantized_fraction": quant_params / total_params, + } + return qmodel, report + + +def main(): + parser = argparse.ArgumentParser() + parser.add_argument("--data-dir", default=os.path.join( + os.path.expanduser("~"), ".cache", "kagglehub", "datasets", "kaka2434", + "wiflow-dataset", "versions", "1", "preprocessed_csi_data")) + parser.add_argument("--subset", type=int, default=10000) + parser.add_argument("--runs-b1", type=int, default=100) + parser.add_argument("--runs-b64", type=int, default=30) + parser.add_argument("--skip-accuracy", action="store_true") + parser.add_argument("--out", default=os.path.join(RESULTS, "edge_optimization.json")) + args = parser.parse_args() + + torch.manual_seed(42) + results = { + "env": { + "torch": torch.__version__, + "platform": platform.platform(), + "processor": platform.processor(), + "num_threads": torch.get_num_threads(), + "checkpoint": os.path.relpath(CHECKPOINT, HERE), + }, + "variants": {}, + } + + # ---- build variants --------------------------------------------------- + fp32 = load_fp32_model() + n_params = sum(p.numel() for p in fp32.parameters()) + results["env"]["params"] = n_params + print(f"fp32 model: {n_params:,} params") + + fp16 = load_fp32_model().half() + + int8, q_report = quantize_int8_dynamic(load_fp32_model()) + results["int8_dynamic_quant_report"] = q_report + print(f"int8 dynamic: {q_report['n_modules_quantized']} modules quantized, " + f"{q_report['params_quantized_fraction']*100:.1f}% of params") + + variants = { + "fp32": (fp32, torch.float32, "retrained_fp32_resaved.pth"), + "fp16": (fp16, torch.float16, "retrained_fp16.pth"), + "int8_dynamic": (int8, torch.float32, "retrained_int8_dynamic.pth"), + } + + # ---- (a) size + (b) latency ------------------------------------------- + for name, (model, dtype, fname) in variants.items(): + path = os.path.join(RESULTS, fname) + size = state_dict_size_bytes(model, path) + print(f"\n=== {name}: {size/1e6:.3f} MB on disk ({fname}) ===") + lat1 = bench_latency(model, 1, args.runs_b1, dtype) + lat64 = bench_latency(model, 64, args.runs_b64, dtype) + print(f" batch 1: {lat1['median_ms_per_window']:.2f} ms/window " + f"({lat1['windows_per_second']:.0f}/s)") + print(f" batch 64: {lat64['median_ms_per_window']:.3f} ms/window " + f"({lat64['windows_per_second']:.0f}/s)") + results["variants"][name] = { + "file": fname, + "size_bytes": size, + "size_mb": size / 1e6, + "latency_batch1": lat1, + "latency_batch64": lat64, + } + + # ---- (c) accuracy ------------------------------------------------------ + if not args.skip_accuracy: + loader, n_clean = build_test_subset(args.data_dir, args.subset) + results["accuracy_subset"] = { + "description": "seed-42 file-level 70/15/15 test split, corrupted " + "windows (files 487-499) excluded, seed-42 random " + "subset", + "subset_size": min(args.subset, n_clean) if args.subset else n_clean, + "clean_test_total": n_clean, + } + for name, (model, dtype, _f) in variants.items(): + print(f"\n=== accuracy: {name} ===") + results["variants"][name]["accuracy"] = evaluate( + model, loader, dtype, label=name) + print(json.dumps(results["variants"][name]["accuracy"], indent=2)) + + # ---- merge into edge_optimization.json --------------------------------- + merged = {} + if os.path.exists(args.out): + with open(args.out) as f: + merged = json.load(f) + merged["torch"] = results + with open(args.out, "w") as f: + json.dump(merged, f, indent=2) + print(f"\nwrote {args.out}") + + +if __name__ == "__main__": + main() diff --git a/benchmarks/wiflow-std/results/edge_optimization.json b/benchmarks/wiflow-std/results/edge_optimization.json new file mode 100644 index 00000000..257d6ad5 --- /dev/null +++ b/benchmarks/wiflow-std/results/edge_optimization.json @@ -0,0 +1,239 @@ +{ + "torch": { + "env": { + "torch": "2.12.0+cpu", + "platform": "Windows-11-10.0.26200-SP0", + "processor": "Intel64 Family 6 Model 197 Stepping 2, GenuineIntel", + "num_threads": 16, + "checkpoint": "results\\retrained_best_pose_model.pth", + "params": 2225042 + }, + "variants": { + "fp32": { + "file": "retrained_fp32_resaved.pth", + "size_bytes": 9068948, + "size_mb": 9.068948, + "latency_batch1": { + "batch_size": 1, + "runs": 100, + "median_ms_per_batch": 24.903650000851485, + "median_ms_per_window": 24.903650000851485, + "windows_per_second": 40.15475642991324 + }, + "latency_batch64": { + "batch_size": 64, + "runs": 30, + "median_ms_per_batch": 184.02919999789447, + "median_ms_per_window": 2.875456249967101, + "windows_per_second": 347.77089723115813 + }, + "accuracy": { + "samples": 10000, + "pck@20": 0.9668200004577636, + "pck@50": 0.9915333324432373, + "mpjpe": 0.00936222033649683, + "wall_seconds": 37.85407733917236 + } + }, + "fp16": { + "file": "retrained_fp16.pth", + "size_bytes": 4580332, + "size_mb": 4.580332, + "latency_batch1": { + "batch_size": 1, + "runs": 100, + "median_ms_per_batch": 23.936699999467237, + "median_ms_per_window": 23.936699999467237, + "windows_per_second": 41.776853117691964 + }, + "latency_batch64": { + "batch_size": 64, + "runs": 30, + "median_ms_per_batch": 102.32584999903338, + "median_ms_per_window": 1.5988414062348966, + "windows_per_second": 625.4529036465817 + }, + "accuracy": { + "samples": 10000, + "pck@20": 0.966773332977295, + "pck@50": 0.9915066654205322, + "mpjpe": 0.009460017587244511, + "wall_seconds": 21.632277250289917 + } + }, + "int8_dynamic": { + "file": "retrained_int8_dynamic.pth", + "size_bytes": 9068948, + "size_mb": 9.068948, + "latency_batch1": { + "batch_size": 1, + "runs": 100, + "median_ms_per_batch": 18.105350000041653, + "median_ms_per_window": 18.105350000041653, + "windows_per_second": 55.23229321707117 + }, + "latency_batch64": { + "batch_size": 64, + "runs": 30, + "median_ms_per_batch": 168.77549999844632, + "median_ms_per_window": 2.6371171874757238, + "windows_per_second": 379.20195763359703 + }, + "accuracy": { + "samples": 10000, + "pck@20": 0.9668200004577636, + "pck@50": 0.9915333324432373, + "mpjpe": 0.00936222033649683, + "wall_seconds": 45.35376596450806 + } + } + }, + "int8_dynamic_quant_report": { + "eligible_module_counts": { + "nn.Linear": 0, + "nn.Conv1d": 21, + "nn.Conv2d": 22 + }, + "modules_actually_quantized": [], + "n_modules_quantized": 0, + "params_total": 2225042, + "params_quantized": 0, + "params_quantized_fraction": 0.0 + }, + "accuracy_subset": { + "description": "seed-42 file-level 70/15/15 test split, corrupted windows (files 487-499) excluded, seed-42 random subset", + "subset_size": 10000, + "clean_test_total": 10000 + } + }, + "onnx": { + "env": { + "torch": "2.12.0+cpu", + "onnxruntime": "1.26.0", + "platform": "Windows-11-10.0.26200-SP0" + }, + "export": { + "mode": "dynamic-batch", + "exporter": "torchscript", + "file": "retrained_fp32_dynamic.onnx", + "size_mb": 8.971781 + }, + "parity": { + "fixture": "results/parity_fixture.npz (batch 2, seed 42)", + "max_abs_diff_vs_stored_fixture": 2.384185791015625e-07, + "max_abs_diff_vs_torch_now": 2.384185791015625e-07, + "pass_lt_1e-4": true + }, + "latency": { + "batch1": { + "batch_size": 1, + "runs": 100, + "median_ms_per_batch": 2.5410999987798277, + "median_ms_per_window": 2.5410999987798277, + "windows_per_second": 393.5303610563043 + }, + "batch64": { + "batch_size": 64, + "runs": 30, + "median_ms_per_batch": 181.95204999938142, + "median_ms_per_window": 2.8430007812403346, + "windows_per_second": 351.7410218803118 + } + }, + "ort_int8_dynamic_supplementary": { + "file": "retrained_int8_ort_dynamic.onnx", + "size_mb": 2.438794, + "runs": true, + "max_abs_diff_vs_fp32_fixture": 0.00827130675315857 + } + }, + "onnx_accuracy": { + "onnx_fp32": { + "samples": 10000, + "pck@20": 0.9668200004577636, + "pck@50": 0.9915333324432373, + "mpjpe": 0.00936222568154335, + "wall_seconds": 22.34790802001953 + }, + "onnx_int8_ort_dynamic": { + "samples": 10000, + "pck@20": 0.965240001964569, + "pck@50": 0.9915466655731201, + "mpjpe": 0.01108054072111845, + "wall_seconds": 55.742953062057495 + } + }, + "latency_controlled_rerun": { + "note": "3 interleaved repetitions per variant, median ms/window; quiet box", + "fp32": { + "batch1_ms_per_window_median": 10.969150001983508, + "batch1_reps": [ + 10.969150001983508, + 12.646450000829645, + 10.49820000116597 + ], + "batch64_ms_per_window_median": 2.2734187500077496, + "batch64_reps": [ + 2.377234374989712, + 2.124126562478068, + 2.2734187500077496 + ] + }, + "fp16": { + "batch1_ms_per_window_median": 24.313550000442774, + "batch1_reps": [ + 25.1078499986761, + 21.856999999727122, + 24.313550000442774 + ], + "batch64_ms_per_window_median": 2.414695312495496, + "batch64_reps": [ + 2.5705156249955508, + 1.7137437499741281, + 2.414695312495496 + ] + }, + "int8_dynamic": { + "batch1_ms_per_window_median": 15.627150000000256, + "batch1_reps": [ + 17.67525000104797, + 14.627999998992891, + 15.627150000000256 + ], + "batch64_ms_per_window_median": 2.0546906250160646, + "batch64_reps": [ + 2.0546906250160646, + 2.03407343752815, + 2.9325796875241394 + ] + }, + "onnx_fp32": { + "batch1_ms_per_window_median": 3.186650001225644, + "batch1_reps": [ + 2.7332500012562377, + 3.1995500012271805, + 3.186650001225644 + ], + "batch64_ms_per_window_median": 1.9893374999924163, + "batch64_reps": [ + 1.5590843750032946, + 1.9893374999924163, + 2.2144343749914697 + ] + }, + "onnx_int8_ort_dynamic": { + "batch1_ms_per_window_median": 6.50984999811044, + "batch1_reps": [ + 6.50984999811044, + 6.455249998907675, + 6.789299999581999 + ], + "batch64_ms_per_window_median": 5.770093750015803, + "batch64_reps": [ + 5.770093750015803, + 3.912374999970325, + 7.8067296875019565 + ] + } + } +} \ No newline at end of file diff --git a/docs/adr/ADR-152-wifi-pose-sota-2026-intake.md b/docs/adr/ADR-152-wifi-pose-sota-2026-intake.md index 1e819d7d..77fe9e35 100644 --- a/docs/adr/ADR-152-wifi-pose-sota-2026-intake.md +++ b/docs/adr/ADR-152-wifi-pose-sota-2026-intake.md @@ -47,7 +47,7 @@ Adopt four changes, ordered by effort-vs-gain: 1. **Record transceiver geometry at enrollment.** `EnrollmentProtocol` gains an optional `NodeGeometry` record per node (position estimate, antenna orientation, inter-node distances where known). Stored alongside the room baseline in the bank; schema-versioned so existing banks remain readable. 2. **Fuse geometry embeddings into specialist training.** Where a specialist head consumes the (future, ADR-150) backbone embedding, concatenate a small learned embedding of `NodeGeometry` — the PerceptAlign mechanism, transplanted to our per-room banks. Statistical specialists (current) ignore it; LoRA heads (ADR-151 P6) consume it. -3. **Adopt the two-checkerboard alignment for the camera-supervised path (ADR-079).** When MediaPipe supervision is used, calibrate camera↔WiFi into one shared 3D frame before regression (<5 min, two checkerboards, a few photos). This is the direct defense against F1 for our 92.9%-PCK@20 pipeline. +3. **Adopt the two-checkerboard alignment for the camera-supervised path (ADR-079).** When MediaPipe supervision is used, calibrate camera↔WiFi into one shared 3D frame before regression (<5 min, two checkerboards, a few photos). This is the direct defense against F1 for our camera-supervised pipeline. ~~92.9%-PCK@20~~ — *that figure was retracted during measurement (b) (2026-06-10): the surviving holdout shows a constant-output model under an absolute (non-torso) threshold on 69 near-static frames; mean predictor scores 100% under the same protocol. The §2.2 no-citation rule now applies to it.* 4. **Evaluate on the PerceptAlign cross-domain dataset** (21 subjects / 7 layouts) as the MERIDIAN cross-layout benchmark — *gated on confirming its license and downloadability* (open question; repo per paper: github.com/Trymore-lab/PerceptAlign). > **Gate resolved (2026-06-10, MEASURED by repo inspection):** repo exists, **MIT license**, dataset downloadable from HuggingFace (5 per-scene repos, raw CSI + separate vision keypoints; Intel 5300, 1TX×3RX×3 ant, 57 subcarriers — same order as ESP32 subcarrier counts; Scene3 ships 3 distinct layouts). Code present, no pretrained weights. Benchmark adoption unblocked; dataset-side license terms inherit HF dataset terms (not separately stated — check at download time). diff --git a/docs/readme-details.md b/docs/readme-details.md index 281498ba..5b09ade0 100644 --- a/docs/readme-details.md +++ b/docs/readme-details.md @@ -50,7 +50,7 @@ See [PR #405](https://github.com/ruvnet/RuView/pull/405) for full details. ### What's New in v0.7.0
-Camera Ground-Truth Training — 92.9% PCK@20 +Camera Ground-Truth Training **v0.7.0 adds camera-supervised pose training** using MediaPipe + real ESP32 CSI data: @@ -76,15 +76,20 @@ node scripts/train-wiflow-supervised.js --data data/paired/*.jsonl --scale lite node scripts/eval-wiflow.js --model models/wiflow-real/wiflow-v1.json --data data/paired/*.jsonl ``` -**Result: 92.9% PCK@20** from a 5-minute data collection session with one ESP32-S3 and one webcam. +> **Accuracy retraction (2026-06-10):** the "92.9% PCK@20" figure previously +> shown here is retracted. A forensic recheck of the surviving eval holdout +> (69 samples) found a constant-output model scored with an absolute +> (non-torso-normalized) threshold on nearly-static frames — a protocol under +> which a trivial mean-pose predictor scores 100%. Torso-normalized PCK@20 on +> the same holdout is ~19% (from that degenerate predictor). No measured +> camera-supervised PCK@20 is currently published (CHANGELOG, PR #535). -| Metric | Before (proxy) | After (camera-supervised) | -|--------|----------------|--------------------------| -| PCK@20 | 0% | **92.9%** | -| Eval loss | 0.700 | **0.082** | -| Bone constraint | N/A | **0.008** | -| Training time | N/A | **19 minutes** | -| Model size | N/A | **974 KB** | +| Metric | Camera-supervised run (protocol retracted) | +|--------|--------------------------------------------| +| Eval loss | 0.082 | +| Bone constraint | 0.008 | +| Training time | 19 minutes | +| Model size | 974 KB | Pre-trained model: [HuggingFace ruv/ruview/wiflow-v1](https://huggingface.co/ruv/ruview) @@ -868,7 +873,7 @@ Download a pre-built binary — no build toolchain needed: | Release | What's included | Tag | |---------|-----------------|-----| -| [v0.7.0](https://github.com/ruvnet/RuView/releases/tag/v0.7.0) | **Latest** — Camera-supervised WiFlow model (92.9% PCK@20), ground-truth training pipeline, ruvector optimizations | `v0.7.0` | +| [v0.7.0](https://github.com/ruvnet/RuView/releases/tag/v0.7.0) | **Latest** — Camera-supervised WiFlow model (accuracy figure retracted 2026-06-10, see above), ground-truth training pipeline, ruvector optimizations | `v0.7.0` | | [v0.6.0](https://github.com/ruvnet/RuView/releases/tag/v0.6.0-esp32) | [Pre-trained models on HuggingFace](https://huggingface.co/ruv/ruview), 17 sensing apps, 51.6% contrastive improvement, 0.008ms inference | `v0.6.0-esp32` | | [v0.5.5](https://github.com/ruvnet/RuView/releases/tag/v0.5.5-esp32) | SNN + MinCut (#348 fix) + CNN spectrogram + WiFlow + multi-freq mesh + graph transformer | `v0.5.5-esp32` | | [v0.5.4](https://github.com/ruvnet/RuView/releases/tag/v0.5.4-esp32) | Cognitum Seed integration ([ADR-069](docs/adr/ADR-069-cognitum-seed-csi-pipeline.md)), 8-dim feature vectors, RVF store, witness chain, security hardening | `v0.5.4-esp32` | diff --git a/docs/user-guide.md b/docs/user-guide.md index 1e149d24..b296ac34 100644 --- a/docs/user-guide.md +++ b/docs/user-guide.md @@ -1747,7 +1747,14 @@ See [ADR-071](adr/ADR-071-ruvllm-training-pipeline.md) and the [pretraining tuto For significantly higher accuracy, use a webcam as a **temporary teacher** during training. The camera captures real 17-keypoint poses via MediaPipe, paired with simultaneous ESP32 CSI data. After training, the camera is no longer needed — the model runs on CSI only. -**Result: 92.9% PCK@20** from a 5-minute collection session. +> **Accuracy note (2026-06-10):** the previously cited "92.9% PCK@20" figure is +> retracted — a forensic recheck of the surviving eval holdout showed it came +> from a constant-output model scored with an absolute (non-torso-normalized) +> threshold on 69 nearly-static frames, a protocol under which a trivial +> mean-pose predictor scores 100%. No measured camera-supervised PCK@20 is +> currently published (see CHANGELOG, PR #535). Treat this workflow as a data +> collection mechanism; accuracy claims will follow a ≥35-minute multi-pose +> collection session evaluated with torso-normalized PCK. ### Requirements