Run models on Replicate from Hugging Face

Use Replicate as an Inference Provider with the standard Hugging Face clients. Just your HF_TOKEN and provider="replicate", no separate integration needed.

Setup

pip install -U huggingface_hub pillow
export HF_TOKEN=hf_...

Create a token at huggingface.co/settings/tokens with the "Make calls to Inference Providers" permission.

Text to image

Model: Tongyi-MAI/Z-Image-Turbo. Also try black-forest-labs/FLUX.1-dev or Qwen/Qwen-Image.

import os
from huggingface_hub import InferenceClient

client = InferenceClient(provider="replicate", api_key=os.environ["HF_TOKEN"])

image = client.text_to_image(
    "A cinematic photo of an astronaut riding a horse",
    model="Tongyi-MAI/Z-Image-Turbo",
)
image.save("astronaut.png")

Image editing (image to image)

Model: black-forest-labs/FLUX.1-Kontext-dev. Also try Qwen/Qwen-Image-Edit or black-forest-labs/FLUX.2-dev.

import os
from huggingface_hub import InferenceClient

client = InferenceClient(provider="replicate", api_key=os.environ["HF_TOKEN"])

image = client.image_to_image(
    "cat.png",
    prompt="Turn the cat into a tiger",
    model="black-forest-labs/FLUX.1-Kontext-dev",
)
image.save("tiger.png")

Text to video

Model: Wan-AI/Wan2.2-T2V-A14B-Diffusers

import os
from huggingface_hub import InferenceClient

client = InferenceClient(provider="replicate", api_key=os.environ["HF_TOKEN"])

video = client.text_to_video(
    "A young man walking on the street at sunset",
    model="Wan-AI/Wan2.2-T2V-A14B-Diffusers",
)
with open("video.mp4", "wb") as f:
    f.write(video)

Speech recognition

Model: openai/whisper-large-v3

import os
from huggingface_hub import InferenceClient

client = InferenceClient(provider="replicate", api_key=os.environ["HF_TOKEN"])

result = client.automatic_speech_recognition(
    "sample.flac",
    model="openai/whisper-large-v3",
)
print(result.text)

Text to speech

Model: ResembleAI/chatterbox

import os
from huggingface_hub import InferenceClient

client = InferenceClient(provider="replicate", api_key=os.environ["HF_TOKEN"])

audio = client.text_to_speech(
    "Hello from Replicate on Hugging Face!",
    model="ResembleAI/chatterbox",
)
with open("speech.wav", "wb") as f:
    f.write(audio)