&&&& RUNNING TensorRT.trtexec [TensorRT v8601] # trtexec --onnx=resnet18.onnx --saveEngine=resnet18-fp32.plan --minShapes=x:1x3x224x224 --optShapes=x:4x3x224x224 --maxShapes=x:16x3x224x224 --memPoolSize=workspace:1024MiB --verbose
[05/07/2024-11:12:11] [I] === Model Options ===
[05/07/2024-11:12:11] [I] Format: ONNX
[05/07/2024-11:12:11] [I] Model: resnet18.onnx
[05/07/2024-11:12:11] [I] Output:
[05/07/2024-11:12:11] [I] === Build Options ===
[05/07/2024-11:12:11] [I] Max batch: explicit batch
[05/07/2024-11:12:11] [I] Memory Pools: workspace: 1024 MiB, dlaSRAM: default, dlaLocalDRAM: default, dlaGlobalDRAM: default
[05/07/2024-11:12:11] [I] minTiming: 1
[05/07/2024-11:12:11] [I] avgTiming: 8
[05/07/2024-11:12:11] [I] Precision: FP32
[05/07/2024-11:12:11] [I] LayerPrecisions: 
[05/07/2024-11:12:11] [I] Layer Device Types: 
[05/07/2024-11:12:11] [I] Calibration: 
[05/07/2024-11:12:11] [I] Refit: Disabled
[05/07/2024-11:12:11] [I] Version Compatible: Disabled
[05/07/2024-11:12:11] [I] TensorRT runtime: full
[05/07/2024-11:12:11] [I] Lean DLL Path: 
[05/07/2024-11:12:11] [I] Tempfile Controls: { in_memory: allow, temporary: allow }
[05/07/2024-11:12:11] [I] Exclude Lean Runtime: Disabled
[05/07/2024-11:12:11] [I] Sparsity: Disabled
[05/07/2024-11:12:11] [I] Safe mode: Disabled
[05/07/2024-11:12:11] [I] Build DLA standalone loadable: Disabled
[05/07/2024-11:12:11] [I] Allow GPU fallback for DLA: Disabled
[05/07/2024-11:12:11] [I] DirectIO mode: Disabled
[05/07/2024-11:12:11] [I] Restricted mode: Disabled
[05/07/2024-11:12:11] [I] Skip inference: Disabled
[05/07/2024-11:12:11] [I] Save engine: resnet18-fp32.plan
[05/07/2024-11:12:11] [I] Load engine: 
[05/07/2024-11:12:11] [I] Profiling verbosity: 0
[05/07/2024-11:12:11] [I] Tactic sources: Using default tactic sources
[05/07/2024-11:12:11] [I] timingCacheMode: local
[05/07/2024-11:12:11] [I] timingCacheFile: 
[05/07/2024-11:12:11] [I] Heuristic: Disabled
[05/07/2024-11:12:11] [I] Preview Features: Use default preview flags.
[05/07/2024-11:12:11] [I] MaxAuxStreams: -1
[05/07/2024-11:12:11] [I] BuilderOptimizationLevel: -1
[05/07/2024-11:12:11] [I] Input(s)s format: fp32:CHW
[05/07/2024-11:12:11] [I] Output(s)s format: fp32:CHW
[05/07/2024-11:12:11] [I] Input build shape: x=1x3x224x224+4x3x224x224+16x3x224x224
[05/07/2024-11:12:11] [I] Input calibration shapes: model
[05/07/2024-11:12:11] [I] === System Options ===
[05/07/2024-11:12:11] [I] Device: 0
[05/07/2024-11:12:11] [I] DLACore: 
[05/07/2024-11:12:11] [I] Plugins:
[05/07/2024-11:12:11] [I] setPluginsToSerialize:
[05/07/2024-11:12:11] [I] dynamicPlugins:
[05/07/2024-11:12:11] [I] ignoreParsedPluginLibs: 0
[05/07/2024-11:12:11] [I] 
[05/07/2024-11:12:11] [I] === Inference Options ===
[05/07/2024-11:12:11] [I] Batch: Explicit
[05/07/2024-11:12:11] [I] Input inference shape: x=4x3x224x224
[05/07/2024-11:12:11] [I] Iterations: 10
[05/07/2024-11:12:11] [I] Duration: 3s (+ 200ms warm up)
[05/07/2024-11:12:11] [I] Sleep time: 0ms
[05/07/2024-11:12:11] [I] Idle time: 0ms
[05/07/2024-11:12:11] [I] Inference Streams: 1
[05/07/2024-11:12:11] [I] ExposeDMA: Disabled
[05/07/2024-11:12:11] [I] Data transfers: Enabled
[05/07/2024-11:12:11] [I] Spin-wait: Disabled
[05/07/2024-11:12:11] [I] Multithreading: Disabled
[05/07/2024-11:12:11] [I] CUDA Graph: Disabled
[05/07/2024-11:12:11] [I] Separate profiling: Disabled
[05/07/2024-11:12:11] [I] Time Deserialize: Disabled
[05/07/2024-11:12:11] [I] Time Refit: Disabled
[05/07/2024-11:12:11] [I] NVTX verbosity: 0
[05/07/2024-11:12:11] [I] Persistent Cache Ratio: 0
[05/07/2024-11:12:11] [I] Inputs:
[05/07/2024-11:12:11] [I] === Reporting Options ===
[05/07/2024-11:12:11] [I] Verbose: Enabled
[05/07/2024-11:12:11] [I] Averages: 10 inferences
[05/07/2024-11:12:11] [I] Percentiles: 90,95,99
[05/07/2024-11:12:11] [I] Dump refittable layers:Disabled
[05/07/2024-11:12:11] [I] Dump output: Disabled
[05/07/2024-11:12:11] [I] Profile: Disabled
[05/07/2024-11:12:11] [I] Export timing to JSON file: 
[05/07/2024-11:12:11] [I] Export output to JSON file: 
[05/07/2024-11:12:11] [I] Export profile to JSON file: 
[05/07/2024-11:12:11] [I] 
[05/07/2024-11:12:12] [I] === Device Information ===
[05/07/2024-11:12:12] [I] Selected Device: NVIDIA GeForce RTX 4090
[05/07/2024-11:12:12] [I] Compute Capability: 8.9
[05/07/2024-11:12:12] [I] SMs: 128
[05/07/2024-11:12:12] [I] Device Global Memory: 24209 MiB
[05/07/2024-11:12:12] [I] Shared Memory per SM: 100 KiB
[05/07/2024-11:12:12] [I] Memory Bus Width: 384 bits (ECC disabled)
[05/07/2024-11:12:12] [I] Application Compute Clock Rate: 2.565 GHz
[05/07/2024-11:12:12] [I] Application Memory Clock Rate: 10.501 GHz
[05/07/2024-11:12:12] [I] 
[05/07/2024-11:12:12] [I] Note: The application clock rates do not reflect the actual clock rates that the GPU is currently running at.
[05/07/2024-11:12:12] [I] 
[05/07/2024-11:12:12] [I] TensorRT version: 8.6.1
[05/07/2024-11:12:12] [I] Loading standard plugins
[05/07/2024-11:12:12] [E] Uncaught exception detected: Unable to open library: libnvinfer_plugin.so.8 due to libcublas.so.12: cannot open shared object file: No such file or directory
&&&& FAILED TensorRT.trtexec [TensorRT v8601] # trtexec --onnx=resnet18.onnx --saveEngine=resnet18-fp32.plan --minShapes=x:1x3x224x224 --optShapes=x:4x3x224x224 --maxShapes=x:16x3x224x224 --memPoolSize=workspace:1024MiB --verbose
