Computer vision · SDK

MediaPipe

Google

Integrate packaged perception tasks into device applications.

Official documentationDownload development spec

GOOGLE MEDIAPIPE / DEVELOPER FIELD GUIDE

Choose an interface. Define a verifiable result.

Company or project context, ten technical entry points and a practical analysis of every reference.

2026-10-08 · Integration not tested

Company / project overview

Original summary of official company or project statements; use the source for the full original page.

Google MediaPipe · About
ProjectMediaPipe is an open-source project within Google AI Edge.
PurposeProvides reusable ML solutions for application developers.
ComponentsTasks APIs and compatible models provide inference; Model Maker supports selected customizations and Studio supports evaluation.
Platform and lifecycleTask and platform availability vary. Use current Solutions guides; legacy solution support differs.
Complete official About page ↗
About reference — TA analysis
AudienceBuyers, partners and product owners shortlisting an AI provider.
Our assessmentUse positioning and company history for initial fit. Obtain project-specific deployment, support and commercial terms separately.

Official project and developer sources reviewed; no provider API was called. Expected outputs below are our illustrative contracts, not measured results.

Top 10 technical entry points

Ten documented technical entry points, selected by task; not ten independent SDKs or an official ranking.

Choose by the work you need to complete
InterfacePurpose and audience
01 · Face DetectorAPI / SDK / developer guideLocate faces and facial key points.Face region UI developers
02 · Face LandmarkerAPI / SDK / developer guideEstimate facial landmarks and supported facial outputs.AR and animation developers
03 · Gesture RecognizerAPI / SDK / developer guideRecognize supported hand gestures.Touchless interface developers
04 · Hand LandmarkerAPI / SDK / developer guideEstimate hand landmarks and handedness.Hand interaction developers
05 · Holistic LandmarkerAPI / SDK / developer guideCombine face, pose and hand landmark estimation.Whole-body interaction teams
06 · Image ClassifierAPI / SDK / developer guideAssign categories using a compatible classification model.Image routing developers
07 · Image EmbedderAPI / SDK / developer guideRepresent images as numeric embeddings.Visual retrieval developers
08 · Image SegmenterAPI / SDK / developer guideSeparate image regions with a compatible segmentation model.Pixel-level UI developers
09 · Object DetectorAPI / SDK / developer guideLocate and categorize model-supported objects.Object-based application teams
10 · Pose LandmarkerAPI / SDK / developer guideEstimate body pose landmarks.Movement visualization teams

Reference analysis

Capabilities summarize official documentation. Constraints and proposed tests are our engineering assessment.

01

API / SDK / developer guide

Face Detector

Face Detector · TA
Target audienceFace region UI developers
Documented capabilityLocate faces and facial key points.
Input → outputImage/frame → face boxes and key points.
Execution environmentSelected Tasks SDK, compatible model and supported platform
Our decision constraintDetection is not identity recognition.
Our suggested validationVerify boxes against authorized annotated frames.
Official reference · Face Detector ↗
02

API / SDK / developer guide

Face Landmarker

Face Landmarker · TA
Target audienceAR and animation developers
Documented capabilityEstimate facial landmarks and supported facial outputs.
Input → outputImage/frame → facial geometry; optional outputs depend on settings.
Execution environmentSelected Tasks SDK, compatible model and supported platform
Our decision constraintGeometry is not evidence of a person’s feelings or intent.
Our suggested validationCheck occlusion and coordinate mapping on target devices.
Official reference · Face Landmarker ↗
03

API / SDK / developer guide

Gesture Recognizer

Gesture Recognizer · TA
Target audienceTouchless interface developers
Documented capabilityRecognize supported hand gestures.
Input → outputImage/frame → gesture categories and hand landmarks.
Execution environmentSelected Tasks SDK, compatible model and supported platform
Our decision constraintOnly configured gesture vocabulary is meaningful.
Our suggested validationTest unknown gestures without triggering a command.
Official reference · Gesture Recognizer ↗
04

API / SDK / developer guide

Hand Landmarker

Hand Landmarker · TA
Target audienceHand interaction developers
Documented capabilityEstimate hand landmarks and handedness.
Input → outputImage/frame → normalized and world landmarks, handedness.
Execution environmentSelected Tasks SDK, compatible model and supported platform
Our decision constraintMirroring and coordinate systems must be explicit.
Our suggested validationCompare mirrored and original input deliberately.
Official reference · Hand Landmarker ↗
05

API / SDK / developer guide

Holistic Landmarker

Holistic Landmarker · TA
Target audienceWhole-body interaction teams
Documented capabilityCombine face, pose and hand landmark estimation.
Input → outputImage/frame → grouped body, face and hand landmarks.
Execution environmentSelected Tasks SDK, compatible model and supported platform
Our decision constraintMissing groups must remain missing rather than zero-filled people.
Our suggested validationVerify partial visibility and group association.
Official reference · Holistic Landmarker ↗
06

API / SDK / developer guide

Image Classifier

Image Classifier · TA
Target audienceImage routing developers
Documented capabilityAssign categories using a compatible classification model.
Input → outputImage → ranked categories and scores.
Execution environmentSelected Tasks SDK, compatible model and supported platform
Our decision constraintCategory vocabulary is determined by the chosen model.
Our suggested validationInclude out-of-domain images and review thresholds.
Official reference · Image Classifier ↗
07

API / SDK / developer guide

Image Embedder

Image Embedder · TA
Target audienceVisual retrieval developers
Documented capabilityRepresent images as numeric embeddings.
Input → outputImage → feature vector.
Execution environmentSelected Tasks SDK, compatible model and supported platform
Our decision constraintVectors are not semantic labels; compare compatible embeddings.
Our suggested validationEvaluate relevant and irrelevant image pairs.
Official reference · Image Embedder ↗
08

API / SDK / developer guide

Image Segmenter

Image Segmenter · TA
Target audiencePixel-level UI developers
Documented capabilitySeparate image regions with a compatible segmentation model.
Input → outputImage → category or confidence masks according to configuration.
Execution environmentSelected Tasks SDK, compatible model and supported platform
Our decision constraintResized masks must be aligned before overlay.
Our suggested validationCheck pixel alignment and known-region overlap.
Official reference · Image Segmenter ↗
09

API / SDK / developer guide

Object Detector

Object Detector · TA
Target audienceObject-based application teams
Documented capabilityLocate and categorize model-supported objects.
Input → outputImage/frame → boxes, categories and scores.
Execution environmentSelected Tasks SDK, compatible model and supported platform
Our decision constraintPretrained classes do not automatically include factory defects.
Our suggested validationMeasure missed objects with domain-specific annotations.
Official reference · Object Detector ↗
10

API / SDK / developer guide

Pose Landmarker

Pose Landmarker · TA
Target audienceMovement visualization teams
Documented capabilityEstimate body pose landmarks.
Input → outputImage/frame → normalized and world pose landmarks.
Execution environmentSelected Tasks SDK, compatible model and supported platform
Our decision constraintPose alone does not establish medical condition or safe operation.
Our suggested validationReview occluded joints and frame-to-frame stability.
Official reference · Pose Landmarker ↗

Official documentation overview

Official documentation overview · Supporting reference analysis
TADevelopers and product owners selecting the supported technical path.
Our validation adviceChoose the operation, supported version and runtime before assigning implementation.
Official reference · Official documentation overview ↗

04 / INPUT · EXPECTED OUTPUT · ACCEPTANCE

Expected results before implementation

Illustrative object-detection contract for a selected Tasks model; other tasks return different structures. No camera or model was run.

Illustrative contract · no API call or model execution · not measured provider output

Object detected

Expected behavior — illustrative, not executed
InputAuthorized image and a compatible object model.
Expected resultok; label, score and box in the declared coordinate system.
AcceptanceModel labels match the UI and boxes align after resizing.

No detection

Expected behavior — illustrative, not executed
InputA valid image with no object above the chosen threshold.
Expected resultok; detections: []; never reuse previous-frame boxes.
AcceptanceEmpty state appears without an error banner.

Model mismatch

Expected behavior — illustrative, not executed
InputA missing or incompatible task model.
Expected resulterror; model initialization failed; data is null.
AcceptanceNo fabricated predictions and a recoverable setup message.

Video sequence

Expected behavior — illustrative, not executed
InputFrames with timestamps in the selected video/live mode.
Expected resultResults are associated with their input frame timestamps.
AcceptanceReject or account for invalid ordering; do not promise a result for every live frame.

Map the expected result to your application

Review the scenarios below and map the actual SDK response into this internal contract. No credential, API call or live output is included. Replace example data only after your own integration test.

Expected output JSON — illustrative, not measured · MediaPipe-expected-results.json
{
  "schema_version": "1.0.0",
  "example": true,
  "provider": "Google MediaPipe",
  "status": "ok",
  "input_id": "sample-001",
  "data": {
    "task": "object_detector",
    "coordinate_system": "pixels-original-image",
    "detections": [
      {
        "label": "sample-object",
        "score": 0.82,
        "box_xywh": [
          12,
          18,
          80,
          60
        ]
      }
    ],
    "model_asset": "record-at-integration"
  },
  "error": null,
  "verification": "illustrative-only;not-provider-response"
}
Download specification, example and acceptance plan
Engineering handoff — our proposed contract and gates
Result statesok means the selected operation returned a valid result; a valid empty result differs from error. Pending work remains pending until a terminal result. Do not infer real-world correctness from API success.
ReproducibilityRecord input ID, library/model version, configuration, runtime and coordinate/score conventions. Examples contain invented sample values.
Acceptance — proposedUse representative authorized samples plus empty, malformed and unavailable-runtime cases. Agree quality and latency targets before running tests. No measured accuracy, cost or speed is claimed.
Operation boundaryIllustrative object-detection contract for a selected Tasks model; other tasks return different structures. No camera or model was run.

Related route to assess: LiteRT

Development assessment

On-device perception benchmark

Development concept · integration not tested

Provider capabilities above are based on official documentation or repositories. The proposed product, inputs, deliverable and acceptance criteria below are our development assessment.

Integration pilot

Validate runtime, schemas or hardware together before estimating a deployable product.

Proposed inputs
A small authorized image dataset, task definition and reference annotations.
Proposed deliverable
A reviewable result with image references, labels or annotation state; keep original files.
Acceptance criterion
Compare the supported operation against a labeled sample; report errors and missing results separately.
Dependencies
Dataset access, image rights and a supported model or annotation project.

Development sequence

  1. Confirm access to MediaPipe, license and the exact supported version.
  2. Prepare the sample above and implement one documented operation for “On-device perception benchmark”.
  3. Normalize the result with source, time and explicit error state; keep the provider response for review.
  4. Run the acceptance criterion before estimating rollout effort or committing a customer deliverable.

How to validate demand

Record product views, documentation clicks, specification downloads and contextual hub clicks. These are event counts, not unique people or completed integrations.

This feasibility assessment uses implementation conditions. No traffic-based rank or delivery-time promise is assigned.

Related products