Welcome to Firefly
Switch language
Firefly Docsss
Last Updated: 2026-07-22 11:25:16

AI Compute Software Stack

Software Stack Architecture

AI Software Stack

Multi-Level Delivery

The SpacemiT AI compute software stack provides multi-level deliverables to address the diverse needs of different users across the AI ecosystem.

End-to-End Model Inference

  • OnnxRuntime

    A SpacemiT inference engine based on ONNX Runtime. By leveraging the SpaceMITExecutionProvider, it delivers optimized inference performance.

  • XSlim

    A model quantization and compression toolchain that supports multiple quantization formats and tuning strategies.

  • Llama.cpp

    A lightweight large-model inference engine that is fully open-source and kept in sync with the upstream community.

  • vLLM

    A popular high-performance framework for large language model inference and serving, supporting native deployment of LLMs.

AI Operator Acceleration Libraries

TBD

AI Programming Languages

  • Triton

    Provides a high-performance AI operator programming experience with a Python-based interface.

Examples

On this page