StableLearn Logo

Search Content

CV 11 min read

SAM3 Tutorial: Meta AI Image Segmentation Model Guide | 4M+ Concepts

Master SAM3: Meta's revolutionary AI segmentation model with 4M+ concepts. Complete tutorial covering installation, Python examples, text prompts for precise image/video segmentation. Boost your CV projects with this open-vocabulary model guide.

Cover image for SAM3 Tutorial: Meta AI Image Segmentation Model Guide | 4M+ Concepts

Published 313 days ago. Content may be outdated.

SAM3 (Segment Anything Model 3) is a revolutionary open-vocabulary segmentation model just released by Meta Superintelligence Labs. Honestly, this model is truly impressive—it not only inherits the powerful capabilities of SAM2 but also achieves genuine concept understanding for the first time.

Simply put, SAM3 can now “understand” what you’re saying! You just need to describe a concept in natural language, and it can find and segment it in images and videos. Plus, this guy has a huge “vocabulary”—it can understand over 4 million different concepts. This “concept segmentation” capability represents another major breakthrough in computer vision.

For example, just say “player wearing white jersey” and SAM3 can instantly find and precisely segment all people matching this description. This ability to “understand human language” is truly impressive!

🚀 Core Advantages

  1. Open-Vocabulary Concept Segmentation: The first model that truly “understands human language” and performs segmentation
  2. Massive Concept Library: Supports 4M+ concepts, 50 times more than existing benchmarks!
  3. Multi-Modal Prompt Support: Whether you want to use text, clicks, or boxes, no problem
  4. Near-Human Performance: Achieves 75-80% of human performance on SA-Co tests, quite impressive
  5. Unified Architecture Design: Detector and tracker work independently without interference, higher efficiency
  6. Existence Judgment: Very practical feature that accurately distinguishes similar concepts (like “red jersey player” vs “white jersey player”)

🏗️ Model Architecture

Core Components

  • Shared Visual Encoder: Efficient visual feature extractor shared by detector and tracker
  • DETR Detector: DETR architecture detector based on text, geometry, and image example conditions
  • SAM2 Tracker: Inherits SAM2’s Transformer encoder-decoder architecture
  • Existence Token: Innovative existence judgment mechanism that improves similar concept distinction
  • Text Encoder: Language understanding module for processing natural language text prompts
  • Multi-Modal Fusion Layer: Cross-modal understanding component that integrates visual and language features

Workflow

StepProcessing StageMain OperationInput TypeOutput Result
1Concept UnderstandingParse text prompts and understand concept semanticsNatural language textConcept embeddings
2Visual Feature ExtractionShared encoder extracts image/video featuresRaw pixel dataHigh-dimensional feature maps
3Cross-Modal FusionFuse visual features and text concept embeddingsVisual features + concept embeddingsMulti-modal representation
4Object DetectionDETR detector locates objects matching conceptsMulti-modal representationBounding boxes + confidence
5Fine SegmentationGenerate high-quality instance segmentation masksDetection results + visual featuresSegmentation masks
6Existence JudgmentExistence token judges if concept truly existsGlobal featuresExistence scores

🛠️ Environment Setup

System Requirements

  • Operating System: Linux, Windows, MacOS
  • Python: 3.12+
  • PyTorch: 2.7+
  • CUDA: 12.6+
  • GPU: NVIDIA GPU (recommended 16GB+ VRAM for 848M parameter model)
  • Memory: 32GB+ RAM (recommended)
  • Storage: At least 50GB available space

1. Create Virtual Environment

   # Create Conda environment (Python 3.12 recommended)
conda create -n sam3 python=3.12
conda deactivate
conda activate sam3

2. Install PyTorch

   # Install PyTorch 2.7 with CUDA 12.6 support
pip install torch==2.7.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cu126

3. Install SAM3

   # Clone project repository
git clone https://github.com/facebookresearch/sam3.git
cd sam3

# Install SAM3
pip install -e .

# Install example notebook dependencies (optional)
pip install -e ".[notebooks]"

# Install development and training dependencies (optional)
pip install -e ".[train,dev]"

4. Get Model Access Permission

Important Note: SAM3 is currently quite “exclusive” and requires permission application:

  1. Visit SAM3 Hugging Face Repository
  2. Apply for access permission and wait for approval
  3. Generate Hugging Face access token
  4. Authenticate:
   # Install huggingface_hub
pip install huggingface_hub

# Login to Hugging Face
hf auth login
# Enter your access token

5. Verify Installation

   # Verify PyTorch and CUDA installation
python -c "import torch; print(f'PyTorch: {torch.__version__}'); print(f'CUDA available: {torch.cuda.is_available()}')"

# Verify SAM3 installation
python -c "from sam3.model_builder import build_sam3_image_model; print('SAM3 installation successful!')"

Tips:

  • SAM3 is a “big guy” (848M parameters), make sure you have enough GPU memory
  • First run will automatically download the model, might take a while
  • If network is slow, try configuring Hugging Face mirror sources

🎯 Quick Start

Image Segmentation Example

Basic Text Prompt Segmentation

   import torch
from PIL import Image
import matplotlib.pyplot as plt
from sam3.model_builder import build_sam3_image_model
from sam3.model.sam3_image_processor import Sam3Processor

# Load SAM3 model
model = build_sam3_image_model()
processor = Sam3Processor(model)

# Load image
image = Image.open("your_image.jpg")
inference_state = processor.set_image(image)

# Here's the magic moment! Use text to describe what you want
text_prompt = "player wearing white jersey"  # You can input any concept
output = processor.set_text_prompt(state=inference_state, prompt=text_prompt)

# Get segmentation results
masks = output["masks"]        # Segmentation masks
boxes = output["boxes"]        # Bounding boxes
scores = output["scores"]      # Confidence scores

# Display results
def show_results(image, masks, boxes, scores, text_prompt):
    fig, axes = plt.subplots(1, len(masks) + 1, figsize=(15, 5))
    
    # Show original image
    axes[0].imshow(image)
    axes[0].set_title("Original")
    axes[0].axis('off')
    
    # Show segmentation results
    for i, (mask, box, score) in enumerate(zip(masks, boxes, scores)):
        axes[i+1].imshow(image)
        axes[i+1].imshow(mask, alpha=0.6, cmap='viridis')
        
        # Draw bounding box
        x1, y1, x2, y2 = box
        rect = plt.Rectangle((x1, y1), x2-x1, y2-y1, 
                           fill=False, color='red', linewidth=2)
        axes[i+1].add_patch(rect)
        
        axes[i+1].set_title(f"Score: {score:.3f}")
        axes[i+1].axis('off')
    
    plt.suptitle(f'Text Prompt: "{text_prompt}"', fontsize=16)
    plt.tight_layout()
    plt.show()

show_results(image, masks, boxes, scores, text_prompt)

Complex Concept Segmentation Example

   # Let's try more complex descriptions to test SAM3's "understanding ability"
complex_prompts = [
    "person wearing glasses",
    "red car",
    "bird flying in the sky",
    "laptop on the table",
    "animal eating grass"
]

for prompt in complex_prompts:
    print(f"\nSegmenting: {prompt}")
    
    # Reset image state
    inference_state = processor.set_image(image)
    
    # Execute segmentation
    output = processor.set_text_prompt(state=inference_state, prompt=prompt)
    
    if len(output["masks"]) > 0:
        print(f"Wow, found {len(output['masks'])} matching objects!")
        # Check best result score
        best_idx = output["scores"].argmax()
        print(f"Best match score: {output['scores'][best_idx]:.3f} (higher is more accurate)")
    else:
        print("Hmm, no matching objects found")

Video Segmentation Example

   import torch
from sam3.model_builder import build_sam3_video_predictor

# Initialize SAM3 video predictor
video_predictor = build_sam3_video_predictor()

# Set video path (can be JPEG folder or MP4 file)
video_path = "your_video.mp4"  # or "./video_frames/" folder

# Start video segmentation session
response = video_predictor.handle_request(
    request=dict(
        type="start_session",
        resource_path=video_path,
    )
)

session_id = response["session_id"]
print(f"Great, video session started, ID: {session_id}")

# Add text prompt at specified frame
frame_index = 0  # Choose a frame as starting frame
text_prompt = "running person"  # Your text prompt

response = video_predictor.handle_request(
    request=dict(
        type="add_prompt",
        session_id=session_id,
        frame_index=frame_index,
        text=text_prompt,
    )
)

# Get segmentation results
output = response["outputs"]
print(f"Found {len(output)} matching objects in frame {frame_index}")

# Process each detected object
for i, obj in enumerate(output):
    object_id = obj["object_id"]
    mask = obj["mask"]
    score = obj["score"]
    
    print(f"Object {i+1}: ID={object_id}, Score={score:.3f}")
    
    # Save mask or perform further processing
    # mask is a numpy array, can be used directly

# Interactive refinement (optional)
# You can add click prompts to fine-tune results
refinement_response = video_predictor.handle_request(
    request=dict(
        type="add_point",
        session_id=session_id,
        frame_index=frame_index,
        object_id=output[0]["object_id"],  # Select first object
        point=[320, 240],  # Click coordinates
        is_positive=True,  # Positive click (foreground)
    )
)

print("Done! Video segmentation and interactive optimization completed!")

Batch Video Processing

   # Process multiple video files
video_files = ["video1.mp4", "video2.mp4", "video3.mp4"]
text_prompts = ["soccer player", "cyclist", "swimmer"]

results = []

for video_file, prompt in zip(video_files, text_prompts):
    print(f"\nProcessing: {video_file} - '{prompt}'")
    
    # Start new session
    response = video_predictor.handle_request(
        request=dict(
            type="start_session",
            resource_path=video_file,
        )
    )
    
    session_id = response["session_id"]
    
    # Add text prompt
    response = video_predictor.handle_request(
        request=dict(
            type="add_prompt",
            session_id=session_id,
            frame_index=0,
            text=prompt,
        )
    )
    
    results.append({
        "video": video_file,
        "prompt": prompt,
        "output": response["outputs"]
    })
    
    print(f"Found {len(response['outputs'])} matching objects")

print(f"\nGreat! Batch processing completed, processed {len(results)} videos in one go")

SAM3 Agent Advanced Features

SAM3 also provides SAM3 Agent for handling more complex text prompts:

   from sam3.agent import Sam3Agent

# Initialize SAM3 Agent
agent = Sam3Agent()

# Load image
image = Image.open("complex_scene.jpg")

# Complex text prompt examples
complex_prompts = [
    "person wearing red shirt and running",
    "blue small car parked on roadside",
    "girl sitting on park bench reading book",
    "white airplane flying in sky",
    "black and white cow eating grass"
]

for prompt in complex_prompts:
    print(f"\nProcessing complex prompt: {prompt}")
    
    # Use Agent to process complex prompts
    result = agent.segment_with_complex_prompt(image, prompt)
    
    if result["success"]:
        masks = result["masks"]
        confidence = result["confidence"]
        print(f"Segmentation successful, confidence: {confidence:.3f}")
        print(f"Found {len(masks)} matching objects")
    else:
        print(f"Segmentation failed: {result['error']}")

🔧 Advanced Features

1. Multi-Concept Simultaneous Segmentation

   # Simultaneously segment multiple different concepts in the same image
from sam3.model.sam3_image_processor import Sam3Processor

processor = Sam3Processor(build_sam3_image_model())
inference_state = processor.set_image(image)

# Define multiple concepts
concepts = [
    "person",
    "car", 
    "building",
    "tree",
    "animal"
]

all_results = {}

for concept in concepts:
    print(f"Segmenting: {concept}")
    
    # Reset state to avoid interference
    inference_state = processor.set_image(image)
    
    # Execute segmentation
    output = processor.set_text_prompt(state=inference_state, prompt=concept)
    
    all_results[concept] = {
        "masks": output["masks"],
        "boxes": output["boxes"],
        "scores": output["scores"]
    }
    
    print(f"Found {len(output['masks'])} {concept} instances")

# Display all results together
def show_multi_concept_results(image, results):
    fig, axes = plt.subplots(2, 3, figsize=(18, 12))
    axes = axes.flatten()
    
    # Show original image
    axes[0].imshow(image)
    axes[0].set_title("Original")
    axes[0].axis('off')
    
    # Show segmentation results for each concept
    for i, (concept, result) in enumerate(results.items()):
        if i >= 5:  # Show maximum 5 concepts
            break
            
        axes[i+1].imshow(image)
        
        # Overlay all masks for this concept
        for mask in result["masks"]:
            axes[i+1].imshow(mask, alpha=0.4, cmap='viridis')
        
        axes[i+1].set_title(f'{concept} ({len(result["masks"])} items)')
        axes[i+1].axis('off')
    
    plt.tight_layout()
    plt.show()

show_multi_concept_results(image, all_results)

2. Fine-Grained Concept Segmentation

   # Use more fine-grained concept descriptions
fine_grained_prompts = [
    "man wearing red T-shirt",
    "woman wearing sunglasses", 
    "black sedan",
    "white small dog",
    "green potted plant",
    "blue umbrella"
]

for prompt in fine_grained_prompts:
    print(f"\nFine-grained segmentation: {prompt}")
    
    inference_state = processor.set_image(image)
    output = processor.set_text_prompt(state=inference_state, prompt=prompt)
    
    if len(output["masks"]) > 0:
        # Get best match
        best_idx = output["scores"].argmax()
        best_score = output["scores"][best_idx]
        
        if best_score > 0.5:  # Set confidence threshold
            print(f"Found high-confidence match: {best_score:.3f}")
            
            # Display results
            plt.figure(figsize=(12, 6))
            
            plt.subplot(1, 2, 1)
            plt.imshow(image)
            plt.title("Original")
            plt.axis('off')
            
            plt.subplot(1, 2, 2)
            plt.imshow(image)
            plt.imshow(output["masks"][best_idx], alpha=0.6, cmap='viridis')
            plt.title(f'{prompt}\nConfidence: {best_score:.3f}')
            plt.axis('off')
            
            plt.tight_layout()
            plt.show()
        else:
            print(f"Low confidence match: {best_score:.3f}")
    else:
        print("No matching objects found")

3. Existence Judgment Feature

   # SAM3's existence judgment feature can accurately determine if concepts exist in images

# Test existence of different concepts
test_concepts = [
    "dog",           # might exist
    "cat",           # might exist  
    "elephant",      # might not exist
    "airplane",      # might not exist
    "car",           # might exist
    "dragon",        # doesn't exist (fictional creature)
]

existence_results = []

for concept in test_concepts:
    inference_state = processor.set_image(image)
    output = processor.set_text_prompt(state=inference_state, prompt=concept)
    
    # Check if there are high-confidence detection results
    if len(output["masks"]) > 0:
        max_score = max(output["scores"])
        exists = max_score > 0.3  # Existence threshold
        
        existence_results.append({
            "concept": concept,
            "exists": exists,
            "confidence": max_score,
            "count": len(output["masks"])
        })
    else:
        existence_results.append({
            "concept": concept,
            "exists": False,
            "confidence": 0.0,
            "count": 0
        })

# Display existence judgment results
print("Concept Existence Analysis:")
print("-" * 50)
for result in existence_results:
    status = "✓ Exists" if result["exists"] else "✗ Not Found"
    print(f"{result['concept']:8} | {status:8} | Confidence: {result['confidence']:.3f} | Count: {result['count']}")

📊 Performance Optimization

1. SAM3 Model Performance Overview

Model Specifications

FeatureSAM3
Parameters848M
ArchitectureDetector + Tracker + Shared Visual Encoder
Supported Concepts4M+ types
Recommended GPU Memory16GB+
Recommended System Memory32GB+

Performance Benchmark Results

Image Segmentation Performance:

DatasetMetricSAM3Human PerformanceBest Comparison Model
SA-Co/GoldcgF154.172.8OWLv2: 24.6
LVIScgF137.2-OWLv2: 29.3
LVISAP48.5-DINO-X: 38.5
COCOAP56.4-DINO-X: 56.0

Video Segmentation Performance:

DatasetMetricSAM3Human Performance
SA-V testcgF130.353.1
SA-V testpHOTA58.070.5
YT-Temporal-1BcgF150.871.2
SmartGlassescgF136.458.5
BURSTHOTA44.5-

Why SAM3 is Impressive:

  • Achieves 75-80% of human performance on SA-Co tests, quite impressive
  • Concept count is 50 times more than existing benchmarks, that’s a huge leap
  • Basically “unrivaled” in the open-vocabulary segmentation field

2. Inference Optimization

Batch Inference Optimization

SAM3 supports batch processing of multiple images to improve efficiency:

   from sam3.model_builder import build_sam3_image_model
from sam3.model.sam3_image_processor import Sam3Processor
import torch

# Initialize model
model = build_sam3_image_model()
processor = Sam3Processor(model)

# Batch processing example
def batch_inference(images, prompts, batch_size=4):
    """
    Batch process multiple images and prompts
    """
    results = []
    
    for i in range(0, len(images), batch_size):
        batch_images = images[i:i+batch_size]
        batch_prompts = prompts[i:i+batch_size]
        
        batch_results = []
        
        with torch.no_grad():  # Disable gradient computation to save memory
            for image, prompt in zip(batch_images, batch_prompts):
                inference_state = processor.set_image(image)
                output = processor.set_text_prompt(
                    state=inference_state, 
                    prompt=prompt
                )
                batch_results.append(output)
        
        results.extend(batch_results)
        
        # Clear GPU memory
        torch.cuda.empty_cache()
    
    return results

# Usage example
images = [Image.open(f"image_{i}.jpg") for i in range(10)]
prompts = ["person", "car", "building", "animal", "plant"] * 2

results = batch_inference(images, prompts, batch_size=2)
print(f"Done! Batch processing completed, processed {len(results)} tasks in one go")

Memory Optimization

   # Memory optimization strategies
import gc

def memory_efficient_inference(image, prompt):
    """
    Memory-optimized inference method
    """
    try:
        # Clear previous memory
        torch.cuda.empty_cache()
        gc.collect()
        
        # Execute inference
        with torch.no_grad():
            inference_state = processor.set_image(image)
            output = processor.set_text_prompt(
                state=inference_state, 
                prompt=prompt
            )
        
        return output
        
    finally:
        # Ensure memory cleanup
        torch.cuda.empty_cache()
        gc.collect()

# Usage example
result = memory_efficient_inference(image, "person wearing red clothes")

Performance Monitoring

   def monitor_performance():
    """
    Monitor SAM3 performance metrics
    """
    import psutil
    import time
    
    # GPU memory usage
    if torch.cuda.is_available():
        gpu_memory = torch.cuda.get_device_properties(0).total_memory
        gpu_allocated = torch.cuda.memory_allocated(0)
        gpu_cached = torch.cuda.memory_reserved(0)
        
        print(f"GPU Total Memory: {gpu_memory / 1e9:.1f} GB")
        print(f"GPU Allocated: {gpu_allocated / 1e9:.1f} GB")
        print(f"GPU Cached: {gpu_cached / 1e9:.1f} GB")
    
    # System memory usage
    memory = psutil.virtual_memory()
    print(f"System Memory Usage: {memory.percent}%")
    print(f"Available Memory: {memory.available / 1e9:.1f} GB")

# Performance testing
def benchmark_sam3(test_images, test_prompts, num_runs=10):
    """
    SAM3 performance benchmark
    """
    import time
    
    times = []
    
    for i in range(num_runs):
        start_time = time.time()
        
        for image, prompt in zip(test_images, test_prompts):
            inference_state = processor.set_image(image)
            output = processor.set_text_prompt(
                state=inference_state, 
                prompt=prompt
            )
        
        end_time = time.time()
        times.append(end_time - start_time)
    
    avg_time = sum(times) / len(times)
    throughput = len(test_images) / avg_time
    
    print(f"Average Processing Time: {avg_time:.3f}s")
    print(f"Processing Speed: {throughput:.1f} images/s")
    
    return avg_time, throughput

🤝 Community and Resources

Learning Resources

SA-Co Dataset

Practical Resources

  • Jupyter Examples:

    • sam3_image_predictor_example.ipynb - Image segmentation and text prompts
    • sam3_video_predictor_example.ipynb - Video segmentation and interactive optimization
    • sam3_image_batched_inference.ipynb - Batch inference examples
    • sam3_agent.ipynb - SAM3 Agent complex prompt processing
  • Hugging Face Models:

Development Resources

  • Code Formatting: ufmt format .
  • Development Environment: pip install -e ".[dev,train]"
  • Contributing Guide: CONTRIBUTING.md

Getting Help

Citing SAM 3

If you use SAM 3 or SA-Co dataset in your research, please cite the relevant papers (BibTeX to be released).

Final Words: SAM3 can be considered Meta Superintelligence Labs’ “masterpiece,” achieving a major breakthrough in the open-vocabulary segmentation field. We recommend keeping an eye on the official repository—who knows when there might be new surprises!

Share Article

More Articles

Related Posts

No related posts yet