Skip to content

Validation status

This page summarizes the evidence shipped with the 0.13.0a1 application source. It is a claim boundary, not a certificate that every node is validated for every assay.

Evidence available now

  • automated tests cover graph behavior, persistence, previews, export, operations, UI behavior, I/O routes, and bundled examples;
  • architecture tests enforce the Qt-free core/ dependency boundary, and focused tests cover stable source revisions, physical grids, detached snapshots, atomic persistence, typed execution, and stale-result rejection;
  • 13 deterministic synthetic samples and 13 checked-in workflows support regression checks and inspection;
  • a generated two-source batch bundle exercises three paired items, nine exact NPY/TIFF/TSV outputs, workflow/config hashes, source identities, manifests, archives, and three finalized item sidecars;
  • an analytical phantom report exercises calibrated object/mesh morphology;
  • method notes and focused tests cover colocalization calculations;
  • tables, CSV/TSV writing, workflow JSON, and Python generation have automated behavior checks;
  • focused UI/execution tests cover isolated-tuning boundaries, actionable and waiting stale states, progressive node previews, compatible layer reuse, and rejection of stale contrast/histogram results;
  • batch tests cover attached-config validation and round-trip, guarded source- axis declarations, representative scientific preflight, restore without preview, direct plan-only execution, complete-item fast skips, transient atomic-write retries, and continuing after a final item-sidecar failure;
  • compute tests cover import-safe CPU-only use, workflow-schema-4 and batch-schema-3 intent, eligibility planning, exact implementation identity, resident device segments, memory admission, classified fallback, optimizer review/apply, Prefer-GPU selection/serialization/UI/durable behavior, progress, cancellation, cleanup, and atomic publication; and
  • opt-in native-Windows RTX tests exercise a real durable GPU batch and an imported generated Python workflow through the same executor.

0.13.0a1 release source and artifacts

The immutable v0.13.0a1 tag resolves to application commit 7520a5bb3ea9fe296bb231c63d1598b833ac10f6 on main. Its exact-source CI run passed package, manifest, lint, distribution-metadata, installed-wheel smoke, and CPython 3.12 and 3.13 tests on Windows, Linux, and macOS. The CI-built distributions are qualification outputs, not publication artifacts.

On native Windows, the exact selected source's complete local suite completed with 4,077 passed, 2 skipped, 2 documented expected failures, 83 warnings, and zero failures. The skips are the opt-in real-CUDA tests and the expected failures are the documented CuPy integer-parity gaps. The exact source was then covered by the cross-platform CI matrix above.

The final tagged artifacts published through PyPI and attached to the GitHub pre-release are:

Artifact SHA-256
napari_vipp-0.13.0a1-py3-none-any.whl C157D2D0E5909A76A1A6093493AB17A245682DF53D83DD0A23FAA0FDF7A3BE02
napari_vipp-0.13.0a1.tar.gz E3231CF22EA2907C3FE05F73477E5E3CD3B10FF2A053BBA92A8045D140E7F7E0

The bundled synthetic-volume example was launched from a clean, non-editable, release-style Windows environment, and the operator explicitly accepted its appearance. A separate headless Prefer-GPU smoke selected cucim-subtract_background-v2 for Subtract Background without fallback, returned a (31, 37) native-uint16 result, and reported clean accelerator cleanup. This is bounded evidence for one example and one declared cuCIM operation region, not broad manual GUI, reader, filesystem, hardware, or assay qualification.

The pinned private Windows cuCIM route was exercised against upstream tag v26.06.00 at commit 3c15781c207eab93a317dd9803a6e726fe01f7c4. Two clean builds from the exact 41-distribution no-dependency lock produced byte-identical wheel files and the policy-pinned canonical payload d640d1e17bcce15d32d03841997252bf915b63da855e406c35f0d70c5a5ea667. Metadata, licenses, exact file inventory, install-helper admission, pip check, and a real GPU import/runtime probe passed. This locally built wheel remains private: VIPP neither ships nor hosts it, and each user must repeat the pinned build for their own environment.

These are the release files. The historical artifacts below must not be uploaded or substituted for them.

The earlier pre-Prefer-GPU candidate at application commit e024409 passed package/manifest, lint, and CPython 3.12 and 3.13 CI on Linux, Windows, and macOS. The candidate wheel and source distribution built and passed Twine metadata checks; a clean CPU-only wheel environment passed installation/import checks, and the installed CUDA wheel environment passed compute-doctor plus the two opt-in real-CUDA durable batch/generated-Python tests. A green CI/package matrix is not equivalent to a manual GUI smoke pass on every operating system, Qt/display environment, filesystem, microscope reader, or GPU. See the candidate CI run. Those results remain evidence for that exact tree, but the Prefer-GPU source change invalidates it as the final release candidate.

The later automated checkpoint is application commit 444f68290fe4359b05c68a027d3ae0a413412fe5 on codex/gpu-cross-platform-support. Its full local suite completed in 299.65 seconds with 3,754 passed, 2 skipped, 2 xfailed, 83 warnings, and zero failures. The skips are the opt-in real-CUDA durable batch smokes. The xfails are the documented CuPy uint8/uint16 integer-parity gaps. Ruff, source and installed-wheel npe2 manifest validation, package build, Twine, and the installed-wheel resource smoke passed. These counts describe that exact historical tree only. Subsequent compute-lifecycle, optimizer, source-loading, generated-CLI, and cleanup-quarantine hardening changed application behavior, so 444f682 is superseded and is not the current release candidate.

The checkpoint's real RTX 5090 Prefer-GPU integration test passed and selected CPU Extract Channel, cuCIM Subtract Background, CPU native-uint16 Gaussian, and CuPyX Median without per-node overrides. The complete pipeline had exact parity, no fallback, and clean accelerator cleanup. This verifies the automated placement contract for that environment; manual Prefer-GPU UI acceptance is still pending.

The exact historical checkpoint artifact hashes are:

  • wheel: 58b08cbb8396c9fe27d28a69d52b06e160817e58987e6dc3b8c5059f3dae9804;
  • source distribution: 00fcb15452d3ed71d344859d4306cbcdcc458686959169a2c47485d8f7abec9b.

These artifacts must not be uploaded for 0.13.0a1. They are superseded by the tagged 7520a5b release source and final hashes above. See the release notes for the historical artifact table.

A bounded source-candidate smoke on 4 August 2026 used application commit e024409 on an Apple M1 Max (arm64) running macOS 26.5.2. Manual checks covered launch/basic CPU processing, collection-batch progress and cooperative cancellation, and memory presentation as one system-RAM budget without a fabricated VRAM total. Four focused tests additionally covered the CPU-safe cancel path, worker-token propagation, single-shot UI cancellation, retained cancelled state, manifest, and cleanup evidence. The follow-up operator record is commit ff21040. This is bounded source-checkout evidence, not a raw test log, clean-wheel Mac qualification, broad reader/display/filesystem coverage, or Apple GPU evidence.

A bounded manual Windows acceptance pass inherited from 0.12.0a3 covered direct unpreviewed batch execution, complete Skip items, continued processing after an item failure, attached-config save/reload, and a representative 3D deconvolution batch. It does not establish cross-platform behavior, broad filesystem interoperability, large-collection scalability, or restoration quality for arbitrary samples.

On 4 August 2026, the operator also reported a bounded native-Windows napari UI smoke pass on the later ff21040 development checkout. A local schema-4 Custom workflow loaded the representative private ND2 acquisition and exercised the intended channel, slice navigation/display behavior, and backend-badge presentation. The retained last-run JSON was subsequently overwritten by an Auto run, so it does not independently preserve the exact mixed-backend assignment or cleanup result. This is operator-attested UI regression evidence only; it is not Prefer-GPU, final-wheel, durable-replay, or broad Windows/GPU qualification.

The e024409 full local suite passed 3,722 tests, with 2 skipped, 2 documented expected failures, and 83 warnings. The two final real-CUDA durable execution tests passed when enabled. These numbers describe that earlier candidate checkout. Preserve them as historical evidence, but do not use them as final-release counts. Use the release-source and final-artifact results above for 0.13.0a1 qualification.

The recorded native-Windows RTX 5090 evidence records operation-level scientific parity, memory, progress/cancellation, cleanup, and timing for the declared public GPU regions. The source-candidate timing refresh measured Richardson-Lucy at 24.898/0.414 seconds CPU/GPU (60.20x) on the private ND2 volume, 35.997/0.455 seconds (79.14x) on the medium synthetic volume, and 137.820/1.517 seconds (90.87x) on the large synthetic volume. Richardson-Lucy TV measured 34.921/0.599 seconds (58.34x) on the private volume and 55.936/0.564 seconds (98.61x) on the medium volume; both RL-TV workloads reported parity and clean cleanup. These are descriptive machine-local results, not portable performance guarantees or durable Auto assignments.

The release's generated calibrated-morphology report records 28/28 checks passed. Its own scope excludes broad numerical equivalence, biological interpretation, and all data conditions. The application test suite checks that the report remains synchronized with the validation script.

Do not generalize this evidence into

  • usability superiority over other tools;
  • broad equivalence to Fiji, CellProfiler, scikit-image, or another package;
  • scalability to whole-slide or high-content datasets;
  • complete OME/acquisition metadata fidelity;
  • biological validity of a segmentation, restoration, or measurement workflow;
  • user-study evidence beyond explicitly described pilot observations.

High-priority evidence gaps

Passing deterministic tests is valuable internal evidence, but it is not the same as an external comparison or assay validation. The distinction matters:

Area Current in-repository evidence Next evidence needed
Watershed/object separation Touching-disk split tests, exported-workflow execution, and 3D-default behavior tests Broader 3D phantoms, split/merge metrics, external comparison, representative real images
Colocalization/association Deterministic metric, overlap, distance, and association tests plus synthetic examples External numerical comparisons and assay-specific positive/negative controls
Skeleton networks Synthetic network workflows and focused operation tests Prespecified topology and calibrated-length packs, perturbation tests, external comparison
I/O and metadata Focused format, dtype, validation, and round-trip tests A release-pinned field matrix and licensed corpus of representative microscope files
PSF/deconvolution Deterministic 2D/3D synthetic images, measured-PSF samples, and operation tests Real bead PSFs, representative microscopy images, artifact/noise analysis, performance characterization
Compute/GPU execution Exact operation-region tests, immutable policy/evidence records, tagged release source 7520a5b with a green package/cross-platform CI matrix and recorded final artifact hashes, a qualified private local cuCIM rebuild, real RTX 5090 Prefer-GPU placement/parity/cleanup and performance checks, OOM/cancellation/cleanup coverage, and bounded M1 Max CPU plus Windows UI operator smokes Qualify native Linux GPU, RTX 40-series Windows, broader clean-host GPU/wheel environments and drivers/runtimes, an Apple provider if pursued, and broader cross-platform manual GUI acceptance
Sources and physical grids Revision-change, owned-snapshot, stale-worker, semantic-axis, scale/unit/origin, mask-broadcast, and image/PSF grid tests Independent corpus covering live readers, network filesystems, registration histories, and heterogeneous microscope metadata
Large data/batch Functional cache/path/memory tests plus deterministic attached/standalone config, planner, direct plan-only execution, source verification, complete-item fast skips, staging, retry, manifest/archive, sidecar, collision, replay, continuation, exact-output bundle, a bounded Windows acceptance pass, and bounded M1 Max CPU progress/cancellation evidence Representative memory/time benchmarks, forced-process interruption studies, large collection stress tests, broader cross-platform/cloud-filesystem studies, semantic-axis iteration, and HCS traversal
Workflow/export architecture Schema-4/schema-3 migration, batch-config/manifest schema 3, guarded source-axis declarations, optional batch-attachment validation, snapshot materialization, atomic-write failure, shared-executor compute provenance, multi-source binding, cancellation, and runtime-version tests Independent reproducibility exercises across archived environments and long-lived release migrations
Usability No release-pinned public usability study Ethics-reviewed, preregistered task study with a controlled comparator and neutral outcomes

Release-specific limitations

  • Workflow schemas 1 and 2 are intentionally rejected. Valid schema-3 workflows load as explicit CPU and save as schema 4; cached pixels/tables are not serialized and exported Python is runtime-version pinned. Recalculate, regenerate exports, and validate after upgrading.
  • Version-1 batch configs load as explicit CPU; version-2 configs retain their saved compute request. Both older versions have no source-axis declaration until reviewed and saved as version 3. A saved Auto, Prefer GPU, or Custom request is intent; actual implementation provenance must be retained from each run. Auto uses reviewed GPU defaults without compatible history; accelerated-only history schedules one same-surface CPU measurement before a later matching run applies the 1.20x/20-ms gate. Prefer GPU instead requests every reviewed eligible accelerator regardless of speed; Custom owns per-node choices and benchmarking.
  • GPU candidates cover only declared operation/dtype/parameter/shape/memory and environment regions. The initial public gate is one exact native-Windows RTX 5090/CUDA 13/CPython 3.12 stack. macOS is CPU-only in this release; the bounded M1 Max CPU smoke above does not admit an Apple accelerator. The CPU/GPU matrix is a readable summary; the runtime policy/decision remains authoritative.
  • cuCIM remains optional and omits Clara I/O. VIPP distributes no Windows wheel; each user builds the exact 26.6.0 tag/commit locally with the pinned recipe. Admission verifies that build's wheel-file hash, the policy-pinned canonical installed payload, source/recipe provenance, and the existing scientific environment/workload gates. This private per-user rebuild is the settled 0.13.0a1 route; VIPP will not host or redistribute the wheel. Clara whole-slide I/O remains outside that build. See the Windows CUDA guide.
  • Batch processing is local-file and sorted-position oriented. It does not iterate selected T/C/Z combinations or discover plate/well/field structure.
  • Many operations are eager even when a source format supports lazy/chunked access; background execution improves responsiveness, not total work.
  • Declared-grid validation cannot prove biological registration or metadata truth. It can only enforce the axes/calibration supplied to VIPP.
  • Richardson-Lucy/TV controls and synthetic tests do not establish a validated restoration parameter range for a real microscope or assay.
  • Manifest and item sidecar writes improve recovery evidence but are not one transaction across all outputs and provenance files.
  • Cooperative progress/cancellation cannot split an opaque library call or file writer into truthful internal percentages.
  • Revised native-intensity colocalization is a source-aligned compatibility implementation targeting Fiji Coloc 2 3.1.0, but independent numerical parity validation is pending. The ImageJ threshold path similarly targets ImageJ 1.54p for scalar uint8, uint16, and float32; Boolean and RGB/RGBA handling are VIPP extensions and are not claimed as ImageJ-exact. Treat both paths as experimental and validate externally before consequential use.

For your workflow

Use validate a workflow to choose assay-specific evidence. If a public claim depends on a gap above, label it as a limitation or produce the required evidence before making the claim.