Edge AI / model serving · SDK

LiteRT

Google

Run machine-learning inference on supported devices.

Official documentationDownload development spec

GOOGLE LITERT / DEVELOPER FIELD GUIDE

Choose an interface. Define a verifiable result.

Company or project context, ten technical entry points and a practical analysis of every reference.

2026-10-08 · Integration not tested

Company / project overview

Original summary of official company or project statements; use the source for the full original page.

Google LiteRT · About
ProjectLiteRT is Google’s on-device model deployment framework, built on TensorFlow Lite.
ScopeCovers model conversion, optimization and runtime execution across supported edge platforms.
Runtime choicesCurrent documentation distinguishes CompiledModel from the compatibility-oriented Interpreter path.
Adjacent productLiteRT-LM adds language-model orchestration; it is a separate integration choice from this tensor inference example.
Complete official About page ↗
About reference — TA analysis
AudienceBuyers, partners and product owners shortlisting an AI provider.
Our assessmentUse positioning and company history for initial fit. Obtain project-specific deployment, support and commercial terms separately.

Official project and developer sources reviewed; no provider API was called. Expected outputs below are our illustrative contracts, not measured results.

Top 10 technical entry points

Ten documented technical entry points, selected by task; not ten independent SDKs or an official ranking.

Choose by the work you need to complete
InterfacePurpose and audience
01 · CompiledModel quickstartAPI / SDK / developer guidePrepare a model and use the on-device runtime.ML application developers
02 · Android integrationAPI / SDK / developer guideSelect an Android runtime integration path.Android application teams
03 · NPU accelerationAPI / SDK / developer guideConfigure supported NPU execution.Hardware optimization teams
04 · Model conversionAPI / SDK / developer guideConvert supported source models for LiteRT.Model export engineers
05 · TensorFlow Lite migrationAPI / SDK / developer guideReview the supported migration paths to LiteRT.Existing TFLite maintainers
06 · Model optimizationAPI / SDK / developer guideReview optimization including post-training quantization.Model efficiency engineers
07 · LiteRT CLI installationAPI / SDK / developer guideInstall command-line model tooling.Build and tooling teams
08 · Microcontroller runtimeAPI / SDK / developer guideRun suitable small models in constrained environments.Embedded firmware teams
09 · Prebuilt C++ SDKAPI / SDK / developer guideIntegrate prebuilt runtime libraries with CMake.Native application developers
10 · Source and release repositoryAPI / SDK / developer guideInspect source, releases and project build entry points.Runtime maintainers

Reference analysis

Capabilities summarize official documentation. Constraints and proposed tests are our engineering assessment.

01

API / SDK / developer guide

CompiledModel quickstart

CompiledModel quickstart · TA
Target audienceML application developers
Documented capabilityPrepare a model and use the on-device runtime.
Input → outputCompatible model and input tensors → output tensors.
Execution environmentSelected LiteRT runtime and supported target hardware
Our decision constraintOutput meaning comes from the model, not the runtime alone.
Our suggested validationVerify shapes, types and model-specific postprocessing.
Official reference · CompiledModel quickstart ↗
02

API / SDK / developer guide

Android integration

Android integration · TA
Target audienceAndroid application teams
Documented capabilitySelect an Android runtime integration path.
Input → outputModel, app and SDK configuration → device inference.
Execution environmentSelected LiteRT runtime and supported target hardware
Our decision constraintAPI availability depends on the selected distribution and device.
Our suggested validationTest lifecycle and memory on the actual minimum target device.
Official reference · Android integration ↗
03

API / SDK / developer guide

NPU acceleration

NPU acceleration · TA
Target audienceHardware optimization teams
Documented capabilityConfigure supported NPU execution.
Input → outputCompatible graph and hardware → accelerated execution.
Execution environmentSelected LiteRT runtime and supported target hardware
Our decision constraintNot every operator or chip supports the same path.
Our suggested validationRecord actual backend and compare numerical results to baseline.
Official reference · NPU acceleration ↗
04

API / SDK / developer guide

Model conversion

Model conversion · TA
Target audienceModel export engineers
Documented capabilityConvert supported source models for LiteRT.
Input → outputSource model → .tflite artifact.
Execution environmentSelected LiteRT runtime and supported target hardware
Our decision constraintA successful export does not prove task quality.
Our suggested validationCompare source and exported model on held-out inputs.
Official reference · Model conversion ↗
05

API / SDK / developer guide

TensorFlow Lite migration

TensorFlow Lite migration · TA
Target audienceExisting TFLite maintainers
Documented capabilityReview the supported migration paths to LiteRT.
Input → outputExisting integration → selected updated runtime path.
Execution environmentSelected LiteRT runtime and supported target hardware
Our decision constraintInterpreter compatibility and CompiledModel adoption are different changes.
Our suggested validationPin old and new dependencies and replay the same fixtures.
Official reference · TensorFlow Lite migration ↗
06

API / SDK / developer guide

Model optimization

Model optimization · TA
Target audienceModel efficiency engineers
Documented capabilityReview optimization including post-training quantization.
Input → outputModel and optional representative data → optimized artifact.
Execution environmentSelected LiteRT runtime and supported target hardware
Our decision constraintQuantization can change accuracy and tensor types.
Our suggested validationMeasure quality drift before accepting any size or speed gain.
Official reference · Model optimization ↗
07

API / SDK / developer guide

LiteRT CLI installation

LiteRT CLI installation · TA
Target audienceBuild and tooling teams
Documented capabilityInstall command-line model tooling.
Input → outputSupported environment and package → available CLI.
Execution environmentSelected LiteRT runtime and supported target hardware
Our decision constraintPin the tool version; installation alone is not deployment.
Our suggested validationRecord version and validate the chosen command on a fixture.
Official reference · LiteRT CLI installation ↗
08

API / SDK / developer guide

Microcontroller runtime

Microcontroller runtime · TA
Target audienceEmbedded firmware teams
Documented capabilityRun suitable small models in constrained environments.
Input → outputSupported model and firmware tensors → inference output.
Execution environmentSelected LiteRT runtime and supported target hardware
Our decision constraintMemory and supported operations differ from desktop runtime.
Our suggested validationCheck tensor memory budget and target-board outputs.
Official reference · Microcontroller runtime ↗
09

API / SDK / developer guide

Prebuilt C++ SDK

Prebuilt C++ SDK · TA
Target audienceNative application developers
Documented capabilityIntegrate prebuilt runtime libraries with CMake.
Input → outputHeaders, libraries and model → native integration.
Execution environmentSelected LiteRT runtime and supported target hardware
Our decision constraintRuntime and accelerator binaries must match platform and version.
Our suggested validationVerify loading and packaged dependencies on a clean target.
Official reference · Prebuilt C++ SDK ↗
10

API / SDK / developer guide

Source and release repository

Source and release repository · TA
Target audienceRuntime maintainers
Documented capabilityInspect source, releases and project build entry points.
Input → outputPinned source revision → reproducible build inputs.
Execution environmentSelected LiteRT runtime and supported target hardware
Our decision constraintRepository head can differ from installed stable packages.
Our suggested validationRecord the exact tag or commit with build options.
Official reference · Source and release repository ↗

Official documentation overview

Official documentation overview · Supporting reference analysis
TADevelopers and product owners selecting the supported technical path.
Our validation adviceChoose the operation, supported version and runtime before assigning implementation.
Official reference · Official documentation overview ↗

04 / INPUT · EXPECTED OUTPUT · ACCEPTANCE

Expected results before implementation

Illustrative tensor inference contract. Tensor shape, dtype, normalization and labels must come from the actual model; no inference or benchmark was run.

Illustrative contract · no API call or model execution · not measured provider output

Valid inference contract

Expected behavior — illustrative, not executed
InputPinned model; tensor shape and type match its signature.
Expected resultok; output tensor metadata and illustrative values.
AcceptanceCompare to a trusted model fixture using agreed tolerances.

Empty application result

Expected behavior — illustrative, not executed
InputA valid inference whose model-specific postprocessor finds no detections.
Expected resultok with an empty detection list; raw tensors still follow the model.
AcceptanceDo not treat an all-zero tensor as universally equivalent to no detections.

Tensor mismatch

Expected behavior — illustrative, not executed
InputWrong dtype, shape or unsupported operation.
Expected resulterror; no usable output; include a bounded diagnostic.
AcceptanceReject the input rather than silently reshaping its meaning.

Accelerator unavailable

Expected behavior — illustrative, not executed
InputRequested GPU/NPU cannot execute the selected model.
Expected resultExplicit failure or a configured CPU fallback with actual backend disclosed.
AcceptanceDo not claim GPU/NPU acceleration or measured latency without evidence.

Map the expected result to your application

Review the scenarios below and map the actual SDK response into this internal contract. No credential, API call or live output is included. Replace example data only after your own integration test.

Expected output JSON — illustrative, not measured · LiteRT-expected-results.json
{
  "schema_version": "1.0.0",
  "example": true,
  "provider": "Google LiteRT",
  "status": "ok",
  "input_id": "sample-001",
  "data": {
    "model_id": "example-model",
    "backend": "example-cpu",
    "outputs": [
      {
        "name": "example_scores",
        "shape": [
          1,
          3
        ],
        "dtype": "float32",
        "values": [
          [
            0.1,
            0.7,
            0.2
          ]
        ]
      }
    ],
    "latency_ms": null
  },
  "error": null,
  "verification": "illustrative-only;not-provider-response"
}
Download specification, example and acceptance plan
Engineering handoff — our proposed contract and gates
Result statesok means the selected operation returned a valid result; a valid empty result differs from error. Pending work remains pending until a terminal result. Do not infer real-world correctness from API success.
ReproducibilityRecord input ID, library/model version, configuration, runtime and coordinate/score conventions. Examples contain invented sample values.
Acceptance — proposedUse representative authorized samples plus empty, malformed and unavailable-runtime cases. Agree quality and latency targets before running tests. No measured accuracy, cost or speed is claimed.
Operation boundaryIllustrative tensor inference contract. Tensor shape, dtype, normalization and labels must come from the actual model; no inference or benchmark was run.

Related route to assess: ONNX Runtime

Development assessment

Mobile inference benchmark

Development concept · integration not tested

Provider capabilities above are based on official documentation or repositories. The proposed product, inputs, deliverable and acceptance criteria below are our development assessment.

Integration pilot

Validate runtime, schemas or hardware together before estimating a deployable product.

Proposed inputs
One pinned model, representative inputs, target runtime and workload.
Proposed deliverable
Predictions or deployment evidence with model version, latency and resource measurements.
Acceptance criterion
Verify outputs against a known sample and measure latency, memory and failed calls on the chosen target.
Dependencies
Model rights, supported operators, memory budget and hardware or serving infrastructure.

Development sequence

  1. Confirm access to LiteRT, license and the exact supported version.
  2. Prepare the sample above and implement one documented operation for “Mobile inference benchmark”.
  3. Normalize the result with source, time and explicit error state; keep the provider response for review.
  4. Run the acceptance criterion before estimating rollout effort or committing a customer deliverable.

How to validate demand

Record product views, documentation clicks, specification downloads and contextual hub clicks. These are event counts, not unique people or completed integrations.

This feasibility assessment uses implementation conditions. No traffic-based rank or delivery-time promise is assigned.

Related products