azure-ai-vision-imageanalysis-py
microsoft/skills
Bilder mit dem Azure AI Vision SDK analysieren: Bildunterschriften und Tags generieren, Objekte erkennen, Text extrahieren (OCR), Personen erkennen und intelligente Ausschnitte vorschlagen.
...Alle erweiternAzure AI Vision SDK zur Bildanalyse für Python
Client-Bibliothek für die Bildanalyse mit Azure AI Vision 4.0, einschließlich Bildunterschriften, Tags, Objekten, OCR und mehr.
Installation
pip install azure-ai-vision-imageanalysis
Umgebungsvariablen
VISION_ENDPOINT=https://.cognitiveservices.azure.com # Erforderlich für alle Authentifizierungsmethoden
AZURE_TOKEN_CREDENTIALS=prod # Nur erforderlich, wenn „DefaultAzureCredential“ in der Produktion verwendet wird
VISION_KEY= # Nur erforderlich für den unten aufgeführten alten API-Schlüssel-Authentifizierungspfad
Authentifizierung und Lebenszyklus
🔑 Für alle folgenden Code-Beispiele gelten zwei Regeln:
- Verwenden Sie vorzugsweise
„DefaultAzureCredential“. Es funktioniert lokal (Azure CLI / VS Code / Developer CLI) und in Azure (verwaltete Identität, Workload-Identität) ohne Codeänderungen. Vermeiden Sie Verbindungszeichenfolgen, Konto-/API-Schlüssel – diese umgehen die Entra-Überwachung und -Rotation.
- Lokale Entwicklung:
„DefaultAzureCredential“funktioniert unverändert.- Produktion: Setzen Sie
AZURE_TOKEN_CREDENTIALS=prod(oderAZURE_TOKEN_CREDENTIALS=), um die Anmeldeinformationskette auf produktionssichere Anmeldeinformationen zu beschränken.- Hüllen Sie jeden Client in einen Kontextmanager, damit HTTP-Transporte, Sockets und Token-Caches deterministisch freigegeben werden:
- Synchron:
mit `(...)` als Client: - Asynchron:
async mitund(...) als Client: async mit DefaultAzureCredential() als Anmeldeinformationen:(ausazure.identity.aio)Code-Schnipsel können diese Konfiguration zwar verkürzen, aber Produktionscode sollte stets beide Regeln befolgen.
import os
from azure.identity import DefaultAzureCredential, ManagedIdentityCredential
from azure.ai.vision.imageanalysis import ImageAnalysisClient
from azure.ai.vision.imageanalysis.models import VisualFeatures
# Lokale Entwicklung: DefaultAzureCredential. Produktion: Setze AZURE_TOKEN_CREDENTIALS=prod oder AZURE_TOKEN_CREDENTIALS=
credential = DefaultAzureCredential(require_envvar=True)
# Oder verwenden Sie in der Produktion direkt spezifische Anmeldedaten:
# Siehe https://learn.microsoft.com/python/api/overview/azure/identity-readme?view=azure-python#credential-classes
# credential = ManagedIdentityCredential()
with ImageAnalysisClient(
endpoint=os.environ["VISION_ENDPOINT"],
credential=credential,
) as client:
result = client.analyze_from_url(
image_url="https://aka.ms/azsdk/image-analysis/sample.jpg",
visual_features=[VisualFeatures.CAPTION],
)
Alte Methode: API-Schlüssel (bestehende schlüsselbasierte Bereitstellungen)
Neuer Code sollte die oben genannte `DefaultAzureCredential` verwenden. Verwenden Sie `AzureKeyCredential` nur, wenn Sie über eine bestehende schlüsselbasierte Bereitstellung verfügen, die noch nicht zu Entra ID migriert wurde – beispielsweise in regulierten Umgebungen, in denen die Einführung von Entra noch nicht abgeschlossen ist.
import os
from azure.core.credentials import AzureKeyCredential
from azure.ai.vision.imageanalysis import ImageAnalysisClient
from azure.ai.vision.imageanalysis.models import VisualFeatures
with ImageAnalysisClient(
endpoint=os.environ["VISION_ENDPOINT"],
credential=AzureKeyCredential(os.environ["VISION_KEY"]),
) as client:
result = client.analyze_from_url(
image_url="https://aka.ms/azsdk/image-analysis/sample.jpg",
visual_features=[VisualFeatures.CAPTION],
)
Bild über URL analysieren
from azure.ai.vision.imageanalysis.models import VisualFeatures
image_url = "https://example.com/image.jpg"
result = client.analyze_from_url(
image_url=image_url,
visual_features=[
VisualFeatures.CAPTION,
VisualFeatures.TAGS,
VisualFeatures.OBJECTS,
VisualFeatures.READ,
VisualFeatures.PEOPLE,
VisualFeatures.SMART_CROPS,
VisualFeatures.DENSE_CAPTIONS
],
gender_neutral_caption=True,
language="en"
)
Bild aus Datei analysieren
with open("image.jpg", "rb") as f:
image_data = f.read()
result = client.analyze(
image_data=image_data,
visual_features=[VisualFeatures.CAPTION, VisualFeatures.TAGS]
)
Bildbeschriftung
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.CAPTION],
gender_neutral_caption=True
)
if result.caption:
print(f"Bildunterschrift: {result.caption.text}")
print(f"Konfidenz: {result.caption.confidence:.2f}")
Dichte Bildunterschriften (mehrere Bereiche)
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.DENSE_CAPTIONS]
)
if result.dense_captions:
for caption in result.dense_captions.list:
print(f"Beschriftung: {caption.text}")
print(f" Konfidenz: {caption.confidence:.2f}")
print(f" Begrenzungsrahmen: {caption.bounding_box}")
Tags
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.TAGS]
)
if result.tags:
for tag in result.tags.list:
print(f"Tag: {tag.name} (Konfidenz: {tag.confidence:.2f})")
Objekterkennung
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.OBJECTS]
)
if result.objects:
for obj in result.objects.list:
print(f"Objekt: {obj.tags[0].name}")
print(f" Konfidenz: {obj.tags[0].confidence:.2f}")
box = obj.bounding_box
print(f" Begrenzungsrahmen: x={box.x}, y={box.y}, w={box.width}, h={box.height}")
OCR (Textextraktion)
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.READ]
)
if result.read:
for block in result.read.blocks:
for line in block.lines:
print(f"Zeile: {line.text}")
print(f" Begrenzungsvieleck: {line.bounding_polygon}")
# Details auf Wortebene
for word in line.words:
print(f" Wort: {word.text} (Konfidenz: {word.confidence:.2f})")
Personenerkennung
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.PEOPLE]
)
if result.people:
for person in result.people.list:
print(f"Person erkannt:")
print(f" Konfidenz: {person.confidence:.2f}")
box = person.bounding_box
print(f" Begrenzungsrahmen: x={box.x}, y={box.y}, w={box.width}, h={box.height}")
Intelligentes Zuschneiden
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.SMART_CROPS],
smart_crops_aspect_ratios=[0,9, 1,33, 1,78] # Hochformat, 4:3, 16:9
)
if result.smart_crops:
for crop in result.smart_crops.list:
print(f"Seitenverhältnis: {crop.aspect_ratio}")
box = crop.bounding_box
print(f" Ausschnittbereich: x={box.x}, y={box.y}, w={box.width}, h={box.height}")
Asyncher Client
from azure.ai.vision.imageanalysis.aio import ImageAnalysisClient
from azure.identity.aio import DefaultAzureCredential
async def analyze_image():
async with DefaultAzureCredential() as credential:
async with ImageAnalysisClient(
endpoint=endpoint,
credential=credential
) as client:
result = await client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.CAPTION]
)
print(result.caption.text)
Visuelle Merkmale
| Merkmal | Beschreibung |
|---|---|
CAPTION |
Ein Satz, der das Bild beschreibt |
DENSE_CAPTIONS |
Bildunterschriften für mehrere Bereiche |
TAGS |
Inhalts-Tags (Objekte, Szenen, Handlungen) |
OBJECTS |
Objekterkennung mit Begrenzungsrahmen |
LESEN |
OCR-Textextraktion |
PERSONEN |
Personenerkennung mit Begrenzungsrahmen |
SMART_CROPS |
Vorgeschlagene Ausschnittbereiche für Miniaturansichten |
Fehlerbehandlung
from azure.core.exceptions import HttpResponseError
try:
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.CAPTION]
)
except HttpResponseError as e:
print(f"Statuscode: {e.status_code}")
print(f"Grund: {e.reason}")
print(f"Meldung: {e.error.message}")
Bildanforderungen
- Formate: JPEG, PNG, GIF, BMP, WEBP, ICO, TIFF, MPO
- Maximale Größe: 20 MB
- Abmessungen: 50 × 50 bis 16.000 × 16.000 Pixel
Bewährte Vorgehensweisen
- Wählen Sie entweder „sync“ ODER „async“ und bleiben Sie dabei. Mischen Sie keine
„azure.ai.vision.imageanalysis“-Sync-Clients mit„azure.ai.vision.imageanalysis.aio“-Async-Clients im selben Aufrufpfad. Wählen Sie pro Modul einen Modus. - Verwenden Sie für Clients und asynchrone Anmeldeinformationen stets Kontextmanager. Umschließen Sie jeden Client
mit `ImageAnalysisClient(...)` als Client:(sync) oderasync mit `ImageAnalysisClient(...)` als Client:(async). Verwenden Sie für „asyncDefaultAzureCredential“aus„azure.identity.aio“ebenfalls„async“ mit „credential:“, damit Tokens und Transportdaten bereinigt werden. - Wählen Sie nur die benötigten Funktionen aus, um Latenz und Kosten zu optimieren
- Verwenden Sie den asynchronen Client für Szenarien mit hohem Durchsatz
- Behandeln Sie `HttpResponseError ` bei ungültigen Bildern oder Authentifizierungsproblemen
- Aktivieren Sie „gender_neutral_caption“ für inklusive Beschreibungen
- Geben Sie die Sprache für lokalisierte Bildunterschriftenan
- Verwenden Sie „smart_crops_aspect_ratios“, das Ihren Anforderungen an Miniaturansichten entspricht
- Zwischenspeichern der Ergebnisse, wenn dasselbe Bild mehrfach analysiert wird
---
name: azure-ai-vision-imageanalysis-py
description: Analyze images using Azure AI Vision SDK: generate captions, tags, detect objects, extract text (OCR), detect people, and suggest smart crops.
license: MIT
---
# Azure AI Vision Image Analysis SDK for Python
Client library for Azure AI Vision 4.0 image analysis including captions, tags, objects, OCR, and more.
## Installation
```bash
pip install azure-ai-vision-imageanalysis
```
## Environment Variables
```bash
VISION_ENDPOINT=https://<resource>.cognitiveservices.azure.com # Required for all auth methods
AZURE_TOKEN_CREDENTIALS=prod # Required only if DefaultAzureCredential is used in production
VISION_KEY=<your-api-key> # Only required for the legacy API-key auth path below
```
## Authentication & Lifecycle
> **🔑 Two rules apply to every code sample below:**
>
> 1. **Prefer `DefaultAzureCredential`.** It works locally (Azure CLI / VS Code / Developer CLI) and in Azure (managed identity, workload identity) with no code change. Avoid connection strings, account/API keys — they bypass Entra audit and rotation.
> - Local dev: `DefaultAzureCredential` works as-is.
> - Production: set `AZURE_TOKEN_CREDENTIALS=prod` (or `AZURE_TOKEN_CREDENTIALS=<specific_credential>`) to constrain the credential chain to production-safe credentials.
> 2. **Wrap every client in a context manager** so HTTP transports, sockets, and token caches are released deterministically:
> - Sync: `with <Client>(...) as client:`
> - Async: `async with <Client>(...) as client:` **and** `async with DefaultAzureCredential() as credential:` (from `azure.identity.aio`)
>
> Snippets may abbreviate this setup, but production code should always follow both rules.
```python
import os
from azure.identity import DefaultAzureCredential, ManagedIdentityCredential
from azure.ai.vision.imageanalysis import ImageAnalysisClient
from azure.ai.vision.imageanalysis.models import VisualFeatures
# Local dev: DefaultAzureCredential. Production: set AZURE_TOKEN_CREDENTIALS=prod or AZURE_TOKEN_CREDENTIALS=<specific_credential>
credential = DefaultAzureCredential(require_envvar=True)
# Or use a specific credential directly in production:
# See https://learn.microsoft.com/python/api/overview/azure/identity-readme?view=azure-python#credential-classes
# credential = ManagedIdentityCredential()
with ImageAnalysisClient(
endpoint=os.environ["VISION_ENDPOINT"],
credential=credential,
) as client:
result = client.analyze_from_url(
image_url="https://aka.ms/azsdk/image-analysis/sample.jpg",
visual_features=[VisualFeatures.CAPTION],
)
```
### Legacy: API Key (existing keyed deployments)
New code should use `DefaultAzureCredential` above. Use `AzureKeyCredential` only if you have an existing keyed deployment that hasn't been migrated to Entra ID yet — for example, regulated environments still completing their Entra rollout.
```python
import os
from azure.core.credentials import AzureKeyCredential
from azure.ai.vision.imageanalysis import ImageAnalysisClient
from azure.ai.vision.imageanalysis.models import VisualFeatures
with ImageAnalysisClient(
endpoint=os.environ["VISION_ENDPOINT"],
credential=AzureKeyCredential(os.environ["VISION_KEY"]),
) as client:
result = client.analyze_from_url(
image_url="https://aka.ms/azsdk/image-analysis/sample.jpg",
visual_features=[VisualFeatures.CAPTION],
)
```
## Analyze Image from URL
```python
from azure.ai.vision.imageanalysis.models import VisualFeatures
image_url = "https://example.com/image.jpg"
result = client.analyze_from_url(
image_url=image_url,
visual_features=[
VisualFeatures.CAPTION,
VisualFeatures.TAGS,
VisualFeatures.OBJECTS,
VisualFeatures.READ,
VisualFeatures.PEOPLE,
VisualFeatures.SMART_CROPS,
VisualFeatures.DENSE_CAPTIONS
],
gender_neutral_caption=True,
language="en"
)
```
## Analyze Image from File
```python
with open("image.jpg", "rb") as f:
image_data = f.read()
result = client.analyze(
image_data=image_data,
visual_features=[VisualFeatures.CAPTION, VisualFeatures.TAGS]
)
```
## Image Caption
```python
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.CAPTION],
gender_neutral_caption=True
)
if result.caption:
print(f"Caption: {result.caption.text}")
print(f"Confidence: {result.caption.confidence:.2f}")
```
## Dense Captions (Multiple Regions)
```python
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.DENSE_CAPTIONS]
)
if result.dense_captions:
for caption in result.dense_captions.list:
print(f"Caption: {caption.text}")
print(f" Confidence: {caption.confidence:.2f}")
print(f" Bounding box: {caption.bounding_box}")
```
## Tags
```python
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.TAGS]
)
if result.tags:
for tag in result.tags.list:
print(f"Tag: {tag.name} (confidence: {tag.confidence:.2f})")
```
## Object Detection
```python
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.OBJECTS]
)
if result.objects:
for obj in result.objects.list:
print(f"Object: {obj.tags[0].name}")
print(f" Confidence: {obj.tags[0].confidence:.2f}")
box = obj.bounding_box
print(f" Bounding box: x={box.x}, y={box.y}, w={box.width}, h={box.height}")
```
## OCR (Text Extraction)
```python
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.READ]
)
if result.read:
for block in result.read.blocks:
for line in block.lines:
print(f"Line: {line.text}")
print(f" Bounding polygon: {line.bounding_polygon}")
# Word-level details
for word in line.words:
print(f" Word: {word.text} (confidence: {word.confidence:.2f})")
```
## People Detection
```python
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.PEOPLE]
)
if result.people:
for person in result.people.list:
print(f"Person detected:")
print(f" Confidence: {person.confidence:.2f}")
box = person.bounding_box
print(f" Bounding box: x={box.x}, y={box.y}, w={box.width}, h={box.height}")
```
## Smart Cropping
```python
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.SMART_CROPS],
smart_crops_aspect_ratios=[0.9, 1.33, 1.78] # Portrait, 4:3, 16:9
)
if result.smart_crops:
for crop in result.smart_crops.list:
print(f"Aspect ratio: {crop.aspect_ratio}")
box = crop.bounding_box
print(f" Crop region: x={box.x}, y={box.y}, w={box.width}, h={box.height}")
```
## Async Client
```python
from azure.ai.vision.imageanalysis.aio import ImageAnalysisClient
from azure.identity.aio import DefaultAzureCredential
async def analyze_image():
async with DefaultAzureCredential() as credential:
async with ImageAnalysisClient(
endpoint=endpoint,
credential=credential
) as client:
result = await client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.CAPTION]
)
print(result.caption.text)
```
## Visual Features
| Feature | Description |
|---------|-------------|
| `CAPTION` | Single sentence describing the image |
| `DENSE_CAPTIONS` | Captions for multiple regions |
| `TAGS` | Content tags (objects, scenes, actions) |
| `OBJECTS` | Object detection with bounding boxes |
| `READ` | OCR text extraction |
| `PEOPLE` | People detection with bounding boxes |
| `SMART_CROPS` | Suggested crop regions for thumbnails |
## Error Handling
```python
from azure.core.exceptions import HttpResponseError
try:
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.CAPTION]
)
except HttpResponseError as e:
print(f"Status code: {e.status_code}")
print(f"Reason: {e.reason}")
print(f"Message: {e.error.message}")
```
## Image Requirements
- Formats: JPEG, PNG, GIF, BMP, WEBP, ICO, TIFF, MPO
- Max size: 20 MB
- Dimensions: 50x50 to 16000x16000 pixels
## Best Practices
1. **Pick sync OR async and stay consistent.** Do not mix `azure.ai.vision.imageanalysis` sync clients with `azure.ai.vision.imageanalysis.aio` async clients in the same call path. Choose one mode per module.
2. **Always use context managers for clients and async credentials.** Wrap every client in `with ImageAnalysisClient(...) as client:` (sync) or `async with ImageAnalysisClient(...) as client:` (async). For async `DefaultAzureCredential` from `azure.identity.aio`, also use `async with credential:` so tokens and transports are cleaned up.
3. **Select only needed features** to optimize latency and cost
4. **Use async client** for high-throughput scenarios
5. **Handle HttpResponseError** for invalid images or auth issues
6. **Enable gender_neutral_caption** for inclusive descriptions
7. **Specify language** for localized captions
8. **Use smart_crops_aspect_ratios** matching your thumbnail requirements
9. **Cache results** when analyzing the same image multiple times
Alle Dateien
0 Dateienazure-ai-vision-imageanalysis-py installieren
Laden Sie die Skill-Dateien herunter und entpacken Sie sie in Ihr Verzeichnis „.claude/skills/“.
ZIP herunterladenKlonen Sie das Repository und kopieren Sie die Skill-Dateien in Ihr Projekt.
git clone https://github.com/microsoft/skills/tree/main/.github/plugins/azure-sdk-python/skills/azure-ai-vision-imageanalysis-py # Copy SKILL.md to your .claude/skills/ directory
Kopieren





Heim
