GPU evaluation & candidate review
Optimize for the device.
Keep the evidence.
Measure each candidate on the actual Jetson Orin. Preserve the model, dataset, evaluator and environment hashes in MLflow, then apply a fixed review policy with the separate DEEPBOM Review companion.
Choose the target first
Record the exact Orin Nano, Orin NX or AGX Orin module, memory, JetPack/L4T, CUDA, TensorRT, power mode and cooling. A workstation measurement does not establish Orin performance. Orin shares system memory between CPU and GPU; available RAM is not dedicated VRAM.
Use an ARM64 runtime compatible with the installed JetPack. For example, JetPack 6.2.2 includes CUDA 12.6, TensorRT 10.3 and cuDNN 9.3. The server's x86 CUDA installation recipe is not a Jetson recipe.
Follow one measured change at a time
- Freeze criteria. Specify representative data, preprocessing, output interpretation, quality limits, p95 latency, memory and power budgets before generating candidates.
- Measure the original. Keep power mode, shapes, batch size, runtime and cooling comparable. Retain warm-up settings, individual timings, runtime placement and thermal observations.
- Simplify the graph. Check constant folding, redundant transposes and casts. Use DEEPBOM to compare structure and preserve public interfaces.
- Evaluate FP16, then INT8. Build TensorRT engines on the target stack. Check numerical error and task quality. Start INT8 with PTQ; consider GPU QAT when the quality budget is not met and training inputs are available.
- Measure the full pipeline. Include decoding, preprocessing, copies and postprocessing. Report isolated GPU timing separately from application latency. Orin Nano has no DLA; other modules require their own DLA placement and fallback evidence.
- Retain each decision. MLflow records measurements; Review binds them to exact artifacts and saves accept, reject or hold against the unchanged policy.
Train with quantization on the GPU
For a PyTorch vision model, use a workstation CUDA GPU and NVIDIA Model Optimizer to calibrate INT8 quantizers and fine-tune the model with simulated quantization. Preserve the model weights and quantizer state. This requires the trainable model, training data and task loss; a deployment ONNX file alone is not that training pipeline.
Export the chosen checkpoint as ONNX with explicit QuantizeLinear/DequantizeLinear nodes. Build and measure the TensorRT engine on the actual Orin. QAT training memory is not reduced to INT8 inference memory, and successful workstation training does not establish Orin performance.
Use separate MLflow training and deployment-evaluation records. DEEPBOM inspects the exported model; Review compares supplied measurements. Neither performs QAT or proves training provenance from Q/DQ nodes. ONNX Runtime CUDA and TensorRT are different execution paths: record and evaluate the backend that will actually ship.
Follow the NVIDIA ModelOpt QAT procedure, including checkpoint restoration with quantizer state. Pin the toolchain and verify that the exported quantization operators and precisions are supported by the selected TensorRT release.
This is a workflow guide, not an executed QAT experiment or an Orin performance result.
Record measurements, then review the candidate
Run your own CUDA or TensorRT evaluator explicitly. Retain the target configuration, representative inputs, warm-up and timing samples, task-quality results and runtime placement observations. Merely enabling a CUDA provider does not prove that every operation executed on the GPU; inspect runtime events and any CPU fallback.
The separate DEEPBOM Review v0.1.0 preview imports externally recorded evidence. It uses DEEPBOM 1.103.0 internally and has its own release cycle. It supports same-format, single-file ONNX with embedded weights or TFLite, up to 128 MiB per model; this is not the full DEEPBOM 2.0.0 format or IR surface.
npm install -g github:JunHwan-Kwon/deepbom-review#v0.1.0
deepbom-review capabilities
Node.js 20+ is required. Follow the evaluation binding and MLflow import instructions to create baseline.receipt.json and candidate.receipt.json. Use actual evaluation-time hashes; missing historical evidence requires re-evaluation or a hold decision. After choosing your own policy.json criteria, run:
deepbom-review prepare \
--baseline baseline.onnx --candidate candidate.onnx \
--policy policy.json \
--baseline-evaluation baseline.receipt.json \
--candidate-evaluation candidate.receipt.json \
--change "Describe the candidate transformation" \
--out request.json
deepbom-review run request.json --out review-001.zip
deepbom-review verify review-001.zip --expected-sha256 SHA256_PRINTED_BY_RUN
Replace the uppercase hash placeholder with the full hash printed by run. Review does not execute the models. Without evaluation receipts, a valid structural review is saved with hold; an internally consistent ZIP is not itself an acceptance decision. TensorRT engine bytes are outside this preview's artifact formats: retain them with their source-model/build bindings in external experiment records.
What the result establishes
| Static evidence | Artifact identity, serialized structure, interface changes, findings and evidence coverage. |
|---|---|
| Observed CUDA placement | Provider events for executed runtime nodes in the profiled inputs. Fused runtime nodes are not a one-to-one count of original ONNX operators. |
| External measurement | Results supplied by an explicitly run evaluator. Hash binding is not an attestation that measurements are truthful. |
| Comparison limits | Review requires identical dataset, evaluator and environment hashes. Changing devices, runtime, power mode or provider requires a separate experiment. |
| Unmeasured outcomes | This guide does not establish application accuracy, successful QAT export, a speedup or Orin performance. Those require observations from the actual experiment. |