Skip to main content
Skip to content

case study

Orchestrating multi-model inference at the edge

A non-client architecture case study for scheduling inference across models, accelerators and intermittent network conditions.

Emil · Published June 28, 2026 · Updated September 22, 2026 · 9 min read

Integrated control and computing systems

Problem

Conventional operating-system abstractions do not directly express inference placement across models with different latency, energy, memory, integrity and connectivity characteristics.

Process

The study reframed scheduling around inference requests and decision constraints, then separated application intent, orchestration, runtime memory and heterogeneous acceleration into explicit layers.

Implementation

The reference architecture defined a cognitive scheduler, multi-objective placement, cache-aware memory management, deterministic fallback under intermittent connectivity and continuous runtime attestation.

Outcomes

The output is a public preprint with reproducible diagrams and a falsifiable architecture. It identifies implementation and comparative benchmarking as open work; no field or client result is claimed.

Non-client technical case study derived from the cited BELTO-affiliated Zenodo preprint. It does not claim a commissioned engagement, operating deployment or commercial result.

Author

Emil Shirokikh

Founder

Founder of Belto Inc. Writes on engineering, venture building and applied intelligence.