Data & Infrastructure

ONNX Runtime

Open Neural Network Exchange runtime — a cross-platform inference engine. Supports many hardware backends. Used for cross-provider portability of models and for on-device deployment where llama.cpp does not fit.

Related terms

Framework dimensions

Next steps

Where this fits

ONNX Runtime is part of the Data & Infrastructure vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.