Skip to main content
Skip to content

research

On-Device Large Language Model Inference at the Network Edge

A technical report on architecture, optimization and cross-platform runtime design for language-model inference on consumer edge hardware.

Emil Shirokikh · Published May 24, 2026 · Updated September 22, 2026 · 12 min read

Edge perception and inference laboratory

Abstract

A technical report, prepared during the Stanford Ignite Program, presenting a framework and architecture for executing large language models directly on consumer hardware at the network edge.

Research question

How can one model artifact execute across heterogeneous consumer devices while respecting memory, thermal, latency and privacy constraints?

Architecture examined

The report describes a cross-platform runtime spanning iOS, Android and web environments, with hardware capability profiling, execution-provider selection and a hybrid retrieval design.

Optimization position

Quantization, runtime selection and memory planning are treated as system concerns rather than isolated model-compression tasks. The paper describes 4-bit post-training quantization and device-aware calibration.

Limits and open questions

This is an architectural technical report, not evidence of a client deployment or commercial result. Production fitness still depends on device populations, sustained thermal behavior, workload distributions, security review and field measurement.

Decision relevance

The work informs whether a private or latency-sensitive AI workload belongs on-device, in a hybrid topology or in centralized infrastructure.

Read the full paper

Independent technical report by Emil Shirokikh, prepared during the Stanford Ignite Program and self-published on Zenodo. The Zenodo record does not list a BELTO affiliation. This is not presented as client work or evidence of a commercial deployment.

Author

Emil Shirokikh

Founder

Founder of Belto Inc. Writes on engineering, venture building and applied intelligence.