Validation status¶
This page summarizes the evidence shipped with the 0.13.0a1 application source. It is a claim boundary, not a certificate that every node is validated for every assay.
Evidence available now¶
- automated tests cover graph behavior, persistence, previews, export, operations, UI behavior, I/O routes, and bundled examples;
- architecture tests enforce the Qt-free
core/dependency boundary, and focused tests cover stable source revisions, physical grids, detached snapshots, atomic persistence, typed execution, and stale-result rejection; - 13 deterministic synthetic samples and 13 checked-in workflows support regression checks and inspection;
- a generated two-source batch bundle exercises three paired items, nine exact NPY/TIFF/TSV outputs, workflow/config hashes, source identities, manifests, archives, and three finalized item sidecars;
- an analytical phantom report exercises calibrated object/mesh morphology;
- method notes and focused tests cover colocalization calculations;
- tables, CSV/TSV writing, workflow JSON, and Python generation have automated behavior checks;
- focused UI/execution tests cover isolated-tuning boundaries, actionable and waiting stale states, progressive node previews, compatible layer reuse, and rejection of stale contrast/histogram results;
- batch tests cover attached-config validation and round-trip, guarded source- axis declarations, representative scientific preflight, restore without preview, direct plan-only execution, complete-item fast skips, transient atomic-write retries, and continuing after a final item-sidecar failure;
- compute tests cover import-safe CPU-only use, workflow-schema-4 and batch-schema-3 intent, eligibility planning, exact implementation identity, resident device segments, memory admission, classified fallback, optimizer review/apply, Prefer-GPU selection/serialization/UI/durable behavior, progress, cancellation, cleanup, and atomic publication; and
- opt-in native-Windows RTX tests exercise a real durable GPU batch and an imported generated Python workflow through the same executor.
0.13.0a1 release source and artifacts¶
The immutable
v0.13.0a1
tag resolves to application commit
7520a5bb3ea9fe296bb231c63d1598b833ac10f6
on main. Its exact-source
CI run
passed package, manifest, lint, distribution-metadata, installed-wheel smoke,
and CPython 3.12 and 3.13 tests on Windows, Linux, and macOS. The CI-built
distributions are qualification outputs, not publication artifacts.
On native Windows, the exact selected source's complete local suite completed with 4,077 passed, 2 skipped, 2 documented expected failures, 83 warnings, and zero failures. The skips are the opt-in real-CUDA tests and the expected failures are the documented CuPy integer-parity gaps. The exact source was then covered by the cross-platform CI matrix above.
The final tagged artifacts published through PyPI and attached to the GitHub pre-release are:
| Artifact | SHA-256 |
|---|---|
napari_vipp-0.13.0a1-py3-none-any.whl |
C157D2D0E5909A76A1A6093493AB17A245682DF53D83DD0A23FAA0FDF7A3BE02 |
napari_vipp-0.13.0a1.tar.gz |
E3231CF22EA2907C3FE05F73477E5E3CD3B10FF2A053BBA92A8045D140E7F7E0 |
The bundled synthetic-volume example was launched from a clean, non-editable,
release-style Windows environment, and the operator explicitly accepted its
appearance. A separate headless Prefer-GPU smoke selected
cucim-subtract_background-v2 for Subtract Background without fallback,
returned a (31, 37) native-uint16 result, and reported clean accelerator
cleanup. This is bounded evidence for one example and one declared cuCIM
operation region, not broad manual GUI, reader, filesystem, hardware, or assay
qualification.
The pinned private Windows cuCIM route was exercised against upstream tag
v26.06.00 at commit
3c15781c207eab93a317dd9803a6e726fe01f7c4. Two clean builds from the exact
41-distribution no-dependency lock produced byte-identical wheel files and the
policy-pinned canonical payload
d640d1e17bcce15d32d03841997252bf915b63da855e406c35f0d70c5a5ea667.
Metadata, licenses, exact file inventory, install-helper admission, pip check,
and a real GPU import/runtime probe passed. This locally built wheel remains
private: VIPP neither ships nor hosts it, and each user must repeat the pinned
build for their own environment.
These are the release files. The historical artifacts below must not be uploaded or substituted for them.
The earlier pre-Prefer-GPU candidate at application commit e024409 passed
package/manifest, lint, and CPython 3.12 and 3.13 CI on Linux, Windows, and
macOS. The candidate
wheel and source distribution built and passed Twine metadata checks; a clean
CPU-only wheel environment passed installation/import checks, and the installed
CUDA wheel environment passed compute-doctor plus the two opt-in real-CUDA
durable batch/generated-Python tests. A green CI/package matrix is not
equivalent to a manual GUI smoke pass on every operating system, Qt/display
environment, filesystem, microscope reader, or GPU. See the
candidate CI run.
Those results remain evidence for that exact tree, but the Prefer-GPU source
change invalidates it as the final release candidate.
The later automated checkpoint is application commit
444f68290fe4359b05c68a027d3ae0a413412fe5
on codex/gpu-cross-platform-support. Its full local suite completed in 299.65
seconds with 3,754 passed, 2 skipped, 2 xfailed, 83 warnings, and zero
failures. The skips are the opt-in real-CUDA durable batch smokes. The xfails
are the documented CuPy uint8/uint16 integer-parity gaps. Ruff, source and
installed-wheel npe2 manifest validation, package build, Twine, and the
installed-wheel resource smoke passed. These counts describe that exact
historical tree only. Subsequent compute-lifecycle, optimizer, source-loading,
generated-CLI, and cleanup-quarantine hardening changed application behavior,
so 444f682 is superseded and is not the current release candidate.
The checkpoint's real RTX 5090 Prefer-GPU integration test passed and selected
CPU Extract Channel, cuCIM Subtract Background, CPU native-uint16 Gaussian,
and CuPyX Median without per-node overrides. The complete pipeline had exact
parity, no fallback, and clean accelerator cleanup. This verifies the automated
placement contract for that environment; manual Prefer-GPU UI acceptance is
still pending.
The exact historical checkpoint artifact hashes are:
- wheel:
58b08cbb8396c9fe27d28a69d52b06e160817e58987e6dc3b8c5059f3dae9804; - source distribution:
00fcb15452d3ed71d344859d4306cbcdcc458686959169a2c47485d8f7abec9b.
These artifacts must not be uploaded for 0.13.0a1. They are superseded by the
tagged 7520a5b release source and final hashes above. See the
release notes
for the historical artifact table.
A bounded source-candidate smoke on 4 August 2026 used application commit
e024409 on an Apple M1 Max (arm64) running macOS 26.5.2. Manual checks
covered launch/basic CPU processing, collection-batch progress and cooperative
cancellation, and memory presentation as one system-RAM budget without a
fabricated VRAM total. Four focused tests additionally covered the
CPU-safe cancel path, worker-token propagation, single-shot UI cancellation,
retained cancelled state, manifest, and cleanup evidence. The follow-up
operator record is commit
ff21040.
This is bounded source-checkout evidence, not a raw test log, clean-wheel Mac
qualification, broad reader/display/filesystem coverage, or Apple GPU evidence.
A bounded manual Windows acceptance pass inherited from 0.12.0a3 covered direct
unpreviewed batch execution, complete Skip items, continued processing after
an item failure, attached-config save/reload, and a representative 3D
deconvolution batch. It does not establish cross-platform behavior, broad
filesystem interoperability, large-collection scalability, or restoration
quality for arbitrary samples.
On 4 August 2026, the operator also reported a bounded native-Windows napari UI
smoke pass on the later ff21040 development checkout. A local schema-4
Custom workflow loaded the representative private ND2 acquisition and
exercised the intended channel, slice navigation/display behavior, and
backend-badge presentation. The retained last-run JSON was subsequently
overwritten by an Auto run, so it does not independently preserve the exact
mixed-backend assignment or cleanup result. This is operator-attested UI
regression evidence only; it is not Prefer-GPU, final-wheel, durable-replay, or
broad Windows/GPU qualification.
The e024409 full local suite passed 3,722 tests, with 2
skipped, 2 documented expected failures, and 83 warnings. The two final
real-CUDA durable execution tests passed when enabled. These numbers describe
that earlier candidate checkout. Preserve them as historical evidence, but do
not use them as final-release counts. Use the release-source and final-artifact
results above for 0.13.0a1 qualification.
The recorded native-Windows RTX 5090 evidence records operation-level scientific parity, memory, progress/cancellation, cleanup, and timing for the declared public GPU regions. The source-candidate timing refresh measured Richardson-Lucy at 24.898/0.414 seconds CPU/GPU (60.20x) on the private ND2 volume, 35.997/0.455 seconds (79.14x) on the medium synthetic volume, and 137.820/1.517 seconds (90.87x) on the large synthetic volume. Richardson-Lucy TV measured 34.921/0.599 seconds (58.34x) on the private volume and 55.936/0.564 seconds (98.61x) on the medium volume; both RL-TV workloads reported parity and clean cleanup. These are descriptive machine-local results, not portable performance guarantees or durable Auto assignments.
The release's generated calibrated-morphology report records 28/28 checks passed. Its own scope excludes broad numerical equivalence, biological interpretation, and all data conditions. The application test suite checks that the report remains synchronized with the validation script.
Do not generalize this evidence into¶
- usability superiority over other tools;
- broad equivalence to Fiji, CellProfiler, scikit-image, or another package;
- scalability to whole-slide or high-content datasets;
- complete OME/acquisition metadata fidelity;
- biological validity of a segmentation, restoration, or measurement workflow;
- user-study evidence beyond explicitly described pilot observations.
High-priority evidence gaps¶
Passing deterministic tests is valuable internal evidence, but it is not the same as an external comparison or assay validation. The distinction matters:
| Area | Current in-repository evidence | Next evidence needed |
|---|---|---|
| Watershed/object separation | Touching-disk split tests, exported-workflow execution, and 3D-default behavior tests | Broader 3D phantoms, split/merge metrics, external comparison, representative real images |
| Colocalization/association | Deterministic metric, overlap, distance, and association tests plus synthetic examples | External numerical comparisons and assay-specific positive/negative controls |
| Skeleton networks | Synthetic network workflows and focused operation tests | Prespecified topology and calibrated-length packs, perturbation tests, external comparison |
| I/O and metadata | Focused format, dtype, validation, and round-trip tests | A release-pinned field matrix and licensed corpus of representative microscope files |
| PSF/deconvolution | Deterministic 2D/3D synthetic images, measured-PSF samples, and operation tests | Real bead PSFs, representative microscopy images, artifact/noise analysis, performance characterization |
| Compute/GPU execution | Exact operation-region tests, immutable policy/evidence records, tagged release source 7520a5b with a green package/cross-platform CI matrix and recorded final artifact hashes, a qualified private local cuCIM rebuild, real RTX 5090 Prefer-GPU placement/parity/cleanup and performance checks, OOM/cancellation/cleanup coverage, and bounded M1 Max CPU plus Windows UI operator smokes |
Qualify native Linux GPU, RTX 40-series Windows, broader clean-host GPU/wheel environments and drivers/runtimes, an Apple provider if pursued, and broader cross-platform manual GUI acceptance |
| Sources and physical grids | Revision-change, owned-snapshot, stale-worker, semantic-axis, scale/unit/origin, mask-broadcast, and image/PSF grid tests | Independent corpus covering live readers, network filesystems, registration histories, and heterogeneous microscope metadata |
| Large data/batch | Functional cache/path/memory tests plus deterministic attached/standalone config, planner, direct plan-only execution, source verification, complete-item fast skips, staging, retry, manifest/archive, sidecar, collision, replay, continuation, exact-output bundle, a bounded Windows acceptance pass, and bounded M1 Max CPU progress/cancellation evidence | Representative memory/time benchmarks, forced-process interruption studies, large collection stress tests, broader cross-platform/cloud-filesystem studies, semantic-axis iteration, and HCS traversal |
| Workflow/export architecture | Schema-4/schema-3 migration, batch-config/manifest schema 3, guarded source-axis declarations, optional batch-attachment validation, snapshot materialization, atomic-write failure, shared-executor compute provenance, multi-source binding, cancellation, and runtime-version tests | Independent reproducibility exercises across archived environments and long-lived release migrations |
| Usability | No release-pinned public usability study | Ethics-reviewed, preregistered task study with a controlled comparator and neutral outcomes |
Release-specific limitations¶
- Workflow schemas 1 and 2 are intentionally rejected. Valid schema-3 workflows load as explicit CPU and save as schema 4; cached pixels/tables are not serialized and exported Python is runtime-version pinned. Recalculate, regenerate exports, and validate after upgrading.
- Version-1 batch configs load as explicit CPU; version-2 configs retain their saved compute request. Both older versions have no source-axis declaration until reviewed and saved as version 3. A saved Auto, Prefer GPU, or Custom request is intent; actual implementation provenance must be retained from each run. Auto uses reviewed GPU defaults without compatible history; accelerated-only history schedules one same-surface CPU measurement before a later matching run applies the 1.20x/20-ms gate. Prefer GPU instead requests every reviewed eligible accelerator regardless of speed; Custom owns per-node choices and benchmarking.
- GPU candidates cover only declared operation/dtype/parameter/shape/memory and environment regions. The initial public gate is one exact native-Windows RTX 5090/CUDA 13/CPython 3.12 stack. macOS is CPU-only in this release; the bounded M1 Max CPU smoke above does not admit an Apple accelerator. The CPU/GPU matrix is a readable summary; the runtime policy/decision remains authoritative.
- cuCIM remains optional and omits Clara I/O. VIPP distributes no Windows wheel; each user builds the exact 26.6.0 tag/commit locally with the pinned recipe. Admission verifies that build's wheel-file hash, the policy-pinned canonical installed payload, source/recipe provenance, and the existing scientific environment/workload gates. This private per-user rebuild is the settled 0.13.0a1 route; VIPP will not host or redistribute the wheel. Clara whole-slide I/O remains outside that build. See the Windows CUDA guide.
- Batch processing is local-file and sorted-position oriented. It does not iterate selected T/C/Z combinations or discover plate/well/field structure.
- Many operations are eager even when a source format supports lazy/chunked access; background execution improves responsiveness, not total work.
- Declared-grid validation cannot prove biological registration or metadata truth. It can only enforce the axes/calibration supplied to VIPP.
- Richardson-Lucy/TV controls and synthetic tests do not establish a validated restoration parameter range for a real microscope or assay.
- Manifest and item sidecar writes improve recovery evidence but are not one transaction across all outputs and provenance files.
- Cooperative progress/cancellation cannot split an opaque library call or file writer into truthful internal percentages.
- Revised native-intensity colocalization is a source-aligned compatibility
implementation targeting Fiji Coloc 2 3.1.0, but independent numerical
parity validation is pending. The ImageJ threshold path similarly targets
ImageJ 1.54p for scalar
uint8,uint16, andfloat32; Boolean and RGB/RGBA handling are VIPP extensions and are not claimed as ImageJ-exact. Treat both paths as experimental and validate externally before consequential use.
For your workflow¶
Use validate a workflow to choose assay-specific evidence. If a public claim depends on a gap above, label it as a limitation or produce the required evidence before making the claim.