Skip to main content
Skip to content

research

Designing the AI-Native Operating System

A reference architecture for orchestrated multi-model inference at the edge, including cognitive scheduling, placement, cache management and runtime attestation.

Emil · Published June 28, 2026 · Updated September 22, 2026 · 14 min read

Edge compute and network infrastructure

Abstract

A preprint proposing an AI-native operating-system architecture centered on multi-model inference scheduling, heterogeneous accelerators, cache management, offline-first fault tolerance and runtime attestation.

Research question

What changes when an operating system schedules probabilistic inference requests across heterogeneous models and accelerators rather than deterministic threads?

Proposed architecture

The paper presents SlyOS as a layered reference architecture: agentic applications, a cognitive scheduling plane, an inference runtime, heterogeneous edge accelerators and optional cloud capacity.

Scheduling model

Inference placement is framed as continuous multi-objective optimization under latency, energy and integrity constraints, including intermittent connectivity and deterministic fallback.

Memory and trust

The work identifies key-value cache management as a central constraint and argues for continuous runtime attestation across heterogeneous inference paths.

Limits and next evidence

This is a preprint and reference architecture, not a claim of a completed client deployment. Comparative benchmarks, implementation evidence and independent review remain appropriate next steps.

Read the full paper

Preprint self-published on Zenodo under CC BY 4.0. Public Zenodo creator metadata explicitly lists the affiliation “Belto.” It is a research architecture, not a client engagement or commercial-outcome claim.

Author

Emil Shirokikh

Founder

Founder of Belto Inc. Writes on engineering, venture building and applied intelligence.