Computer vision · API

Google Cloud Vision API

Google Cloud

Extract image labels and text through a managed vision service.

Official documentationDownload development spec

GOOGLE CLOUD VISION / DEVELOPER FIELD GUIDE

Choose an interface. Define a verifiable result.

Company or project context, ten technical entry points and a practical analysis of every reference.

2026-10-08 · Integration not tested

Company / project overview

Original summary of official company or project statements; use the source for the full original page.

Google Cloud Vision · About
MissionGoogle describes its mission as making information easier for people to access and use.
Technology ecosystemCompany information connects Google Cloud, Research, DeepMind, Labs and developer resources.
Company contextThe official page links company history, locations, careers and Alphabet investor information.
CommitmentsIt presents accessibility, sustainability, social impact, education and public-policy initiatives.
Procurement decisionUse the company overview for context; evaluate the specific Vision API terms, architecture and acceptance results before choosing it.
Complete official About page ↗
About reference — TA analysis
AudienceBuyers, partners and product owners shortlisting an AI provider.
Our assessmentUse positioning and company history for initial fit. Obtain project-specific deployment, support and commercial terms separately.

Official company and developer pages read on 2026-10-08. Capabilities are source summaries; constraints and test plans are our engineering assessment. No provider call was made.

Top 10 technical entry points

Ten distinct integration entry points selected for industrial catalogues, document intake and image triage; not an official ranking or ten separate products.

Choose by the work you need to complete
InterfacePurpose and audience
01 · Python client librarySDKCall Vision through the official google-cloud-vision package and ADC authentication.Python / backend integrator
02 · images.annotateREST APISubmit image annotation requests over HTTP through the v1 interface.Non-Python backend / API platform
03 · Label DetectionImage classification APIReturns general image labels in English with scores.Catalogue editor / image-triage developer
04 · Text and document OCRText extraction APITEXT_DETECTION extracts image text; DOCUMENT_TEXT_DETECTION returns document-oriented text structure.Document ingestion / technical content team
05 · Object LocalizationBounding-region APIDetects multiple objects with names, scores and normalized bounding polygons.Inspection UI / warehouse prototype team
06 · Asynchronous PDF / TIFF OCRDocument batch APIReads PDFs or TIFFs from Cloud Storage and writes OCR JSON to a destination bucket.Catalogue migration / document pipeline
07 · Asynchronous image batchBulk annotation APIRuns image annotation asynchronously and stores response files in Cloud Storage.Data engineering / large image archive
08 · SafeSearchContent-screening APIReturns likelihood categories for adult, spoof, medical, violence and racy content.Public upload moderation / product owner
09 · Web DetectionImage discovery APIFinds web entities and matching or visually similar image references.Catalogue provenance / content research team
10 · Image PropertiesImage attribute APIReturns general image attributes such as dominant colours.Catalogue visual QA / media developer

Reference analysis

Capabilities summarize official documentation. Constraints and proposed tests are our engineering assessment.

01

SDK

Python client library

Python client library · TA
Target audiencePython / backend integrator
Documented capabilityCall Vision through the official google-cloud-vision package and ADC authentication.
Input → outputImage request → typed annotation response
Execution environmentPython backend connects to Google Cloud; installing the SDK does not put the model on your device.
Our decision constraintSeparate SDK version, credential identity and billing project in configuration.
Our suggested validationPin the resolved package version and verify a permitted single-image call before batching.
Official reference · Python client library ↗
02

REST API

images.annotate

images.annotate · TA
Target audienceNon-Python backend / API platform
Documented capabilitySubmit image annotation requests over HTTP through the v1 interface.
Input → outputImage and selected features → annotation batch response
Execution environmentAuthenticated backend HTTPS request.
Our decision constraintHTTP success is not sufficient: inspect each image response for an error.
Our suggested validationInject a mixed valid/invalid batch and preserve the source-image/result correspondence.
Official reference · images.annotate ↗
03

Image classification API

Label Detection

Label Detection · TA
Target audienceCatalogue editor / image-triage developer
Documented capabilityReturns general image labels in English with scores.
Input → outputImage → labels, confidence and optional entity identifiers
Execution environmentManaged cloud inference.
Our decision constraintGeneral labels do not establish defect detection or identify exact industrial part numbers.
Our suggested validationCompare against a human-labelled sample; keep original English labels when adding translations.
Official reference · Label Detection ↗
04

Text extraction API

Text and document OCR

Text and document OCR · TA
Target audienceDocument ingestion / technical content team
Documented capabilityTEXT_DETECTION extracts image text; DOCUMENT_TEXT_DETECTION returns document-oriented text structure.
Input → outputImage → text and layout information
Execution environmentCloud OCR; documented US/EU processing options are feature-specific.
Our decision constraintOCR is not a bill-of-materials parser or proof of drawing tolerance accuracy.
Our suggested validationMeasure character errors and critical-field exact matches using rotated, low-contrast and multilingual samples.
Official reference · Text and document OCR ↗
05

Bounding-region API

Object Localization

Object Localization · TA
Target audienceInspection UI / warehouse prototype team
Documented capabilityDetects multiple objects with names, scores and normalized bounding polygons.
Input → outputImage → object regions
Execution environmentManaged cloud inference with client-side overlay rendering.
Our decision constraintNormalize coordinates against the actual displayed image; do not assume support for a custom defect class.
Our suggested validationCheck overlays after resize and rotation; measure missed objects and duplicate detections.
Official reference · Object Localization ↗
06

Document batch API

Asynchronous PDF / TIFF OCR

Asynchronous PDF / TIFF OCR · TA
Target audienceCatalogue migration / document pipeline
Documented capabilityReads PDFs or TIFFs from Cloud Storage and writes OCR JSON to a destination bucket.
Input → outputStored document → operation status and paged JSON output
Execution environmentCloud Storage plus asynchronous Vision operation.
Our decision constraintRequires input read/output write permission; reusing an output path can overwrite prior results.
Our suggested validationUse a unique job prefix and reconcile expected page numbers with output and per-page errors.
Official reference · Asynchronous PDF / TIFF OCR ↗
07

Bulk annotation API

Asynchronous image batch

Asynchronous image batch · TA
Target audienceData engineering / large image archive
Documented capabilityRuns image annotation asynchronously and stores response files in Cloud Storage.
Input → outputImage batch → long-running operation and JSON outputs
Execution environmentManaged cloud jobs and storage.
Our decision constraintAccepted submission is not completion; retain the operation identity before a worker restarts.
Our suggested validationResume polling an existing job, inspect final errors and reconcile every submitted image once.
Official reference · Asynchronous image batch ↗
09

Image discovery API

Web Detection

Web Detection · TA
Target audienceCatalogue provenance / content research team
Documented capabilityFinds web entities and matching or visually similar image references.
Input → outputImage → entities and web/image URLs
Execution environmentCloud analysis and external web references.
Our decision constraintSimilarity does not prove ownership, supplier identity or permission to reuse an image.
Our suggested validationManually verify candidate sources; handle expired URLs and never republish matches automatically.
Official reference · Web Detection ↗
10

Image attribute API

Image Properties

Image Properties · TA
Target audienceCatalogue visual QA / media developer
Documented capabilityReturns general image attributes such as dominant colours.
Input → outputImage → colour attributes
Execution environmentManaged cloud analysis.
Our decision constraintLighting, background and white balance confound manufacturing colour checks.
Our suggested validationUse controlled reference photos; do not equate a dominant-colour match with material conformity.
Official reference · Image Properties ↗

ADC and authentication

ADC and authentication · Supporting reference analysis
TABackend / platform owner: configure identity and quota-project permissions.
Our validation adviceLocal gcloud login and application-default login are different. Verify the effective identity and Service Usage Consumer permission; prefer attached or federated identities for deployment.
Official reference · ADC and authentication ↗

Python ImageAnnotatorClient reference

Python ImageAnnotatorClient reference · Supporting reference analysis
TABackend / QA: verify exact request, timeout and retry arguments.
Our validation adviceThe starter uses batch_annotate_images with one image, retry=None and timeout=30. An RPC deadline does not bound credential discovery or the whole process.
Official reference · Python ImageAnnotatorClient reference ↗

Vision pricing

Vision pricing · Supporting reference analysis
TAProduct owner / FinOps: budget by image/page and selected feature, with separate storage costs.
Our validation adviceRecord actual account usage and currency before estimating volume. Multiple features can add billable units; no price or throughput benchmark is measured here.
Official reference · Vision pricing ↗

Quotas and limits

Quotas and limits · Supporting reference analysis
TAPlatform engineer: distinguish request limits, feature quotas and asynchronous capacity.
Our validation adviceCheck current project quotas; bound queue size and concurrency, and exercise throttling without uncontrolled retries.
Official reference · Quotas and limits ↗

Data Usage FAQ

Data Usage FAQ · Supporting reference analysis
TAData owner: distinguish online processing, asynchronous retention and your own stored outputs.
Our validation adviceReview the current FAQ and account terms. Define retention for your own input/output buckets and logs; the provider FAQ does not delete your storage for you.
Official reference · Data Usage FAQ ↗

Vision documentation overview

Vision documentation overview · Supporting reference analysis
TAPM / developer: navigate feature guides, quickstarts and API references.
Our validation adviceChoose a documented operation and confirm whether the task needs a pretrained feature or a separate custom model.
Official reference · Vision documentation overview ↗

04 / INPUT · EXPECTED OUTPUT · ACCEPTANCE

A concrete first integration

One local image sent to cloud LABEL_DETECTION. Generic labels only; OCR, boxes, video and custom defects need separate adapters and validation.

Starter v1.0.0 · offline fixture checks only · provider integration not run

Run it in your own environment

Use an isolated Python environment; install google-cloud-vision and google-auth, then freeze the resolved versions. Select a Google Cloud project with billing and Vision API enabled. For local development run gcloud auth application-default login, then gcloud auth application-default set-quota-project YOUR_PROJECT_ID; your identity needs serviceusage.services.use on that project. Set GOOGLE_CLOUD_QUOTA_PROJECT and IMAGE_PATH. Save as google_vision_smoke.py and run python google_vision_smoke.py. For deployed workloads use an authorized attached/federated identity. The SDK uses ADC; an API key is not this example’s authentication method.

Keep credentials in your local or server-side credential provider. Running this sample sends an image to provider cloud and may consume account usage.

View runnable Python starter · google_vision_smoke.py
"""Smart Tools: one cloud label request. Offline-tested; live integration untested."""
import json
import math
import os
import sys
from datetime import datetime, timezone
from importlib.metadata import version
from pathlib import Path
from time import monotonic


class ProviderFailure(Exception):
    pass


def normalize(raw):
    if not isinstance(raw, dict) or type(raw.get('error_code')) is not int:
        raise ValueError('unexpected_schema')
    if raw['error_code'] != 0:
        raise ProviderFailure()
    labels = raw.get('labels')
    if not isinstance(labels, list):
        raise ValueError('unexpected_schema')
    result = []
    for item in labels:
        if not isinstance(item, dict):
            raise ValueError('unexpected_schema')
        label, score = item.get('label'), item.get('score')
        if not isinstance(label, str) or not label.strip():
            raise ValueError('unexpected_schema')
        if type(score) not in (int, float) or not math.isfinite(score) or not 0 <= score <= 1:
            raise ValueError('unexpected_schema')
        result.append({'label': label, 'score': score})
    return result


def run():
    start = monotonic()
    result = {'schema_version': '1.0.0', 'provider': 'Google Cloud Vision',
              'status': 'error', 'model_version': None, 'feature': 'LABEL_DETECTION',
              'endpoint': 'vision.googleapis.com', 'data': None, 'error': None,
              'captured_at': datetime.now(timezone.utc).isoformat()}
    try:
        project = os.environ['GOOGLE_CLOUD_QUOTA_PROJECT'].strip()
        if not project:
            raise ValueError('configuration')
        with Path(os.environ['IMAGE_PATH']).open('rb') as image:
            content = image.read(5 * 1024 * 1024 + 1)
        # Local 5 MiB policy for this starter; not a quoted provider maximum.
        if not content or len(content) > 5 * 1024 * 1024:
            raise ValueError('starter_image_limit')
        import google.auth
        from google.auth.exceptions import DefaultCredentialsError, RefreshError
        from google.api_core import exceptions as api_errors
        from google.cloud import vision_v1
        result['sdk_version'] = version('google-cloud-vision')
        try:
            credentials, _ = google.auth.default(
                scopes=['https://www.googleapis.com/auth/cloud-platform'],
                quota_project_id=project)
            request = {'requests': [{'image': {'content': content},
                       'features': [{'type_': vision_v1.Feature.Type.LABEL_DETECTION,
                                     'max_results': 10}]}]}
            # Bytes leave this machine. No automatic RPC retry; cloud model revision is not exposed here.
            with vision_v1.ImageAnnotatorClient(
                    credentials=credentials,
                    client_options={'api_endpoint': 'vision.googleapis.com'}) as client:
                batch = client.batch_annotate_images(request=request, retry=None, timeout=30)
            if len(batch.responses) != 1:
                raise ValueError('unexpected_schema')
            response = batch.responses[0]
            raw = {'error_code': response.error.code,
                   'labels': [{'label': x.description, 'score': x.score}
                              for x in response.label_annotations]}
            result['data'] = normalize(raw)
            result['status'] = 'ok'
        except (DefaultCredentialsError, RefreshError):
            result['error'] = 'authentication'
        except (api_errors.Unauthenticated, api_errors.PermissionDenied):
            result['error'] = 'authentication_or_permission'
        except api_errors.ResourceExhausted:
            result['error'] = 'quota_or_capacity'
        except api_errors.DeadlineExceeded:
            result['error'] = 'deadline_exceeded'
        except api_errors.InvalidArgument:
            result['error'] = 'invalid_request'
        except api_errors.GoogleAPICallError:
            result['error'] = 'provider_or_transport_failure'
    except (KeyError, OSError):
        result['error'] = 'configuration'
    except ImportError:
        result['error'] = 'dependency_missing'
    except ProviderFailure:
        result['error'] = 'provider_image_error'
    except (ValueError, TypeError, AttributeError):
        result['error'] = 'configuration_or_schema'
    except Exception:
        # Never print provider exception strings, credential paths or image contents.
        result['error'] = 'unexpected_failure'
    result['elapsed_ms'] = round((monotonic() - start) * 1000)
    return result


if __name__ == '__main__':
    output = run()
    print(json.dumps(output, ensure_ascii=False, allow_nan=False))
    sys.exit(0 if output['status'] == 'ok' else 1)
Download specification, example and acceptance plan
Engineering handoff — our proposed contract and gates
Output contractstatus is ok/error. data is a validated label/score array (empty is valid) or null on failure. error is a redacted category. Scores remain on 0–1; model_version is null because the returned response does not expose a pinned model revision.
Authentication / configurationKeep credentials server-side; do not commit ADC files or service-account keys. Verify quota project, API enablement and billing independently. The sample caps local input at 5 MiB; this is our policy, not a provider limit.
Failure and time budgetNo automatic RPC retries; timeout=30 bounds the RPC, not credential discovery or the entire process. Use a supervised process deadline. Distinguish credentials/permission, quota, deadline, invalid request and per-image errors before deciding whether to retry.
Acceptance gate — proposedPrepare 20 authorised representative images and a human-labelled baseline. Measure precision at an agreed threshold and p95 latency; test blank/corrupt input, denied permission and timeout. Require zero silent failures. A general-object score cannot certify defect-free production.
Cost / data pathImage bytes are sent to Google Cloud. Count selected features per image/page and include surrounding storage/compute costs. Record sample ID, SDK version, feature, time and endpoint without logging image bytes or credentials. No live cost, latency or quality benchmark is available yet.

Related route to assess: Amazon Rekognition

Development assessment

Equipment label reader

Development concept · integration not tested

Provider capabilities above are based on official documentation or repositories. The proposed product, inputs, deliverable and acceptance criteria below are our development assessment.

Small prototype

Test one documented operation with a small real sample after access is confirmed. This is a development judgment, not a delivery estimate.

Proposed inputs
A small authorized image dataset, task definition and reference annotations.
Proposed deliverable
A reviewable result with image references, labels or annotation state; keep original files.
Acceptance criterion
Compare the supported operation against a labeled sample; report errors and missing results separately.
Dependencies
Dataset access, image rights and a supported model or annotation project.

Development sequence

  1. Confirm access to Google Cloud Vision API, license and the exact supported version.
  2. Prepare the sample above and implement one documented operation for “Equipment label reader”.
  3. Normalize the result with source, time and explicit error state; keep the provider response for review.
  4. Run the acceptance criterion before estimating rollout effort or committing a customer deliverable.

How to validate demand

Record product views, documentation clicks, specification downloads and contextual hub clicks. These are event counts, not unique people or completed integrations.

This feasibility assessment uses implementation conditions. No traffic-based rank or delivery-time promise is assigned.

Related products