第 2 课
图片与视频理解(image_url / video_url)
用 image_url / video_url 对 step-5-preview 发图片与视频理解请求,并对照官方格式、数量与大小限制。
学习位置仅保存在当前浏览器,有效期 180 天。
图文讲义
图片与视频理解(image_url / video_url)
接上一课:已能文本对话后,本课按官方 Quickstart「Image understanding / Video understanding」与模型页限制表,用
step-5-preview发出图、视频多模态请求。
主要来源:Quickstart · Step 5 Preview · Pricing
你将得到什么
- 可复制的
image_url/video_urlPython 与 cURL 示例(模型 ID 固定为step-5-preview) - 官方支持的图片/视频格式与数量、大小建议
- 如何把文档里的
example.com占位换成你自己可公网访问的 URL
课前准备
- 上一课的密钥与环境仍可用(
YOUR_STEP_API_KEY或STEPFUN_API_KEY)。 - 准备至少一张公网可访问的图片 URL,以及(可选)一条视频 URL。
重要: 官方示例里的https://example.com/image.jpg/https://example.com/video.mp4只是占位,直接请求通常会失败。请换成你自己托管、对象存储或可直链下载的地址,并确保 StepFun 服务端能拉取到该 URL(需公网可达,注意鉴权与防盗链)。 - 先对照官方限制,避免一上来就超限。
一、官方输入限制(对照表)
摘自 Step 5 Preview「Image and video input」:
| 项 | 支持范围 |
|---|---|
| 图片输入方式 | URL 或 Base64 |
| 图片格式 | JPG/JPEG、PNG、WebP、静态 GIF |
| 单次请求图片数 | 最多 60 张 |
| 图片 detail | low、high |
| 视频输入方式 | URL、Base64,或 Files API 的 stepfile:// 引用 |
| 视频格式 | MP4、QuickTime、Matroska |
| 视频 URL 限制 | 单个 MP4 小于 128 MB;建议时长 小于 5 分钟 |
step-5-preview 与 step-3.7-flash 支持图/视频;step-3.5-flash 不支持图/视频输入。本课始终使用 step-5-preview。
更细的最佳实践见官方「Image understanding best practices」「Video understanding best practices」(Quickstart 文末链接)。
二、动手:图片理解(image_url)
把下面代码里的图片 URL 换成你自己的可达地址,再运行。
Python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_STEP_API_KEY",
base_url="https://api.stepfun.ai/v1",
)
response = client.chat.completions.create(
model="step-5-preview",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "Extract the key information from this image and summarize the main conclusions."},
{
"type": "image_url",
# Replace this placeholder with the URL of your own accessible image.
"image_url": {"url": "https://example.com/image.jpg"},
},
],
}],
)
print(response.choices[0].message.content)
cURL
# Replace this placeholder with the URL of your own accessible image.
curl https://api.stepfun.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_STEP_API_KEY" \
-d '{
"model": "step-5-preview",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Extract the key information from this image and summarize the main conclusions."},
{
"type": "image_url",
"image_url": {"url": "https://example.com/image.jpg"}
}
]
}]
}'
如何替换占位 URL(跟做步骤):
- 把一张 JPG/PNG/WebP/静态 GIF 上传到你能控制的公开空间(对象存储公开读、个人站点、临时图床等)。
- 浏览器无痕窗口打开该 URL,确认能直接看到图片(不是登录页)。
- 把示例中的
https://example.com/image.jpg整段替换成你的 URL,保存后再跑。 - 需要更高细节时,可在
image_url对象里加"detail": "high"(官方支持low/high;默认行为以 API 参考为准)。
对照结果(成功时):
- HTTP 200 / Python 无异常
choices[0].message.content是对图片内容的文字摘要或要点(提示词要求「Extract the key information…」)- 若返回拉取失败类错误:先本机
curl -I 你的图片URL看是否 200、Content-Type 是否为图片;再检查是否防盗链/需 Cookie
常见踩坑:
- 用了需要登录才能看的网盘分享页 → 改成直链
- 动态 GIF / 非列出格式 → 换成 JPG/PNG/WebP/静态 GIF
- 一次塞超过 60 张 → 拆请求
三、动手:视频理解(video_url)
同样,把 https://example.com/video.mp4 换成你自己可达的 MP4(或官方支持的 QuickTime / Matroska)直链。建议单个文件 < 128 MB,时长优先 < 5 分钟。
Python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_STEP_API_KEY",
base_url="https://api.stepfun.ai/v1",
)
response = client.chat.completions.create(
model="step-5-preview",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "Summarize the main content of this video in chronological order."},
{
"type": "video_url",
# Replace this placeholder with the URL of your own accessible video.
"video_url": {"url": "https://example.com/video.mp4"},
},
],
}],
)
print(response.choices[0].message.content)
cURL
# Replace this placeholder with the URL of your own accessible video.
curl https://api.stepfun.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_STEP_API_KEY" \
-d '{
"model": "step-5-preview",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Summarize the main content of this video in chronological order."},
{
"type": "video_url",
"video_url": {"url": "https://example.com/video.mp4"}
}
]
}]
}'
对照结果(成功时):
- 返回文本按时间顺序概括视频主要内容(提示词为 chronological summary)
- 正文仍在
choices[0].message.content - 大文件或长视频可能更慢:先用短样例冒烟,再换正式素材
进阶输入(本课不展开代码,知悉即可):
- 除 URL 外,官方还支持视频 Base64,以及 Files API 上传后的
stepfile://引用——适合不便公开直链的场景。详见模型页与 Video best practices。
四、读响应与计费提醒
非流式请求的读法与上一课相同:
response.choices[0].message.content
JSON 形状回顾:
{
"model": "step-5-preview",
"choices": [
{
"message": {
"role": "assistant",
"content": "The model's response"
},
"finish_reason": "stop"
}
]
}
图片/视频会按 token 计入输入(与文本一样计费)。当前 step-5-preview 标价仍为:输入 cache-miss $1.00 / cache-hit $0.10 / 输出 $2.70(每 1M tokens)。多模态请求通常比纯短文本更费 token,建议先小图、短视频验证链路,再放大素材。
五、本课自检清单
-
model仍是step-5-preview -
content是数组,同时包含type: text与type: image_url或type: video_url - 已替换所有
example.com占位为你自己的可达 URL - 图片格式在 JPG/PNG/WebP/静态 GIF 内,单次 ≤ 60 张
- 视频为 MP4/QuickTime/Matroska;MP4 URL 尽量 < 128 MB、< 5 分钟
- 成功时仍从
choices[0].message.content取正文
六、本课未覆盖(需要时再查官方)
- 流式输出、Tool calling、JSON Mode / JSON Schema
reasoning_effort(low/medium/high)与 Messages API 的output_config.effort- Prompt caching 细则、Claude Code / Step Plan 集成
- 账号档位与 RPM/TPM 限流表(见 Pricing「Tiered Rate Limits」)
做完两课,你已经走通:密钥 → 文本 → 图/视频。上线前请再对照官方 API Reference 补超时、重试、错误处理与密钥管理。
