第 2 课

图片与视频理解(image_url / video_url)

用 image_url / video_url 对 step-5-preview 发图片与视频理解请求,并对照官方格式、数量与大小限制。

图文25 分钟

课程目录第 2 / 2 课

学习位置仅保存在当前浏览器,有效期 180 天。

01 / 图文教材

图文讲义

图片与视频理解(image_url / video_url)

接上一课:已能文本对话后,本课按官方 Quickstart「Image understanding / Video understanding」与模型页限制表,用 step-5-preview 发出图、视频多模态请求。
主要来源:Quickstart · Step 5 Preview · Pricing

你将得到什么

  • 可复制的 image_url / video_url Python 与 cURL 示例(模型 ID 固定为 step-5-preview)
  • 官方支持的图片/视频格式与数量、大小建议
  • 如何把文档里的 example.com 占位换成你自己可公网访问的 URL

课前准备

  1. 上一课的密钥与环境仍可用(YOUR_STEP_API_KEY 或 STEPFUN_API_KEY)。
  2. 准备至少一张公网可访问的图片 URL,以及(可选)一条视频 URL。
    重要: 官方示例里的 https://example.com/image.jpg / https://example.com/video.mp4 只是占位,直接请求通常会失败。请换成你自己托管、对象存储或可直链下载的地址,并确保 StepFun 服务端能拉取到该 URL(需公网可达,注意鉴权与防盗链)。
  3. 先对照官方限制,避免一上来就超限。

一、官方输入限制(对照表)

摘自 Step 5 Preview「Image and video input」:

项支持范围
图片输入方式URL 或 Base64
图片格式JPG/JPEG、PNG、WebP、静态 GIF
单次请求图片数最多 60 张
图片 detaillow、high
视频输入方式URL、Base64,或 Files API 的 stepfile:// 引用
视频格式MP4、QuickTime、Matroska
视频 URL 限制单个 MP4 小于 128 MB;建议时长 小于 5 分钟

step-5-preview 与 step-3.7-flash 支持图/视频;step-3.5-flash 不支持图/视频输入。本课始终使用 step-5-preview。

更细的最佳实践见官方「Image understanding best practices」「Video understanding best practices」(Quickstart 文末链接)。

二、动手:图片理解(image_url)

把下面代码里的图片 URL 换成你自己的可达地址,再运行。

Python

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_STEP_API_KEY",
    base_url="https://api.stepfun.ai/v1",
)

response = client.chat.completions.create(
    model="step-5-preview",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Extract the key information from this image and summarize the main conclusions."},
            {
                "type": "image_url",
                # Replace this placeholder with the URL of your own accessible image.
                "image_url": {"url": "https://example.com/image.jpg"},
            },
        ],
    }],
)

print(response.choices[0].message.content)

cURL

# Replace this placeholder with the URL of your own accessible image.
curl https://api.stepfun.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_STEP_API_KEY" \
  -d '{
    "model": "step-5-preview",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "Extract the key information from this image and summarize the main conclusions."},
        {
          "type": "image_url",
          "image_url": {"url": "https://example.com/image.jpg"}
        }
      ]
    }]
  }'

如何替换占位 URL(跟做步骤):

  1. 把一张 JPG/PNG/WebP/静态 GIF 上传到你能控制的公开空间(对象存储公开读、个人站点、临时图床等)。
  2. 浏览器无痕窗口打开该 URL,确认能直接看到图片(不是登录页)。
  3. 把示例中的 https://example.com/image.jpg 整段替换成你的 URL,保存后再跑。
  4. 需要更高细节时,可在 image_url 对象里加 "detail": "high"(官方支持 low / high;默认行为以 API 参考为准)。

对照结果(成功时):

  • HTTP 200 / Python 无异常
  • choices[0].message.content 是对图片内容的文字摘要或要点(提示词要求「Extract the key information…」)
  • 若返回拉取失败类错误:先本机 curl -I 你的图片URL 看是否 200、Content-Type 是否为图片;再检查是否防盗链/需 Cookie

常见踩坑:

  • 用了需要登录才能看的网盘分享页 → 改成直链
  • 动态 GIF / 非列出格式 → 换成 JPG/PNG/WebP/静态 GIF
  • 一次塞超过 60 张 → 拆请求

三、动手:视频理解(video_url)

同样,把 https://example.com/video.mp4 换成你自己可达的 MP4(或官方支持的 QuickTime / Matroska)直链。建议单个文件 < 128 MB,时长优先 < 5 分钟。

Python

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_STEP_API_KEY",
    base_url="https://api.stepfun.ai/v1",
)

response = client.chat.completions.create(
    model="step-5-preview",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Summarize the main content of this video in chronological order."},
            {
                "type": "video_url",
                # Replace this placeholder with the URL of your own accessible video.
                "video_url": {"url": "https://example.com/video.mp4"},
            },
        ],
    }],
)

print(response.choices[0].message.content)

cURL

# Replace this placeholder with the URL of your own accessible video.
curl https://api.stepfun.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_STEP_API_KEY" \
  -d '{
    "model": "step-5-preview",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "Summarize the main content of this video in chronological order."},
        {
          "type": "video_url",
          "video_url": {"url": "https://example.com/video.mp4"}
        }
      ]
    }]
  }'

对照结果(成功时):

  • 返回文本按时间顺序概括视频主要内容(提示词为 chronological summary)
  • 正文仍在 choices[0].message.content
  • 大文件或长视频可能更慢:先用短样例冒烟,再换正式素材

进阶输入(本课不展开代码,知悉即可):

  • 除 URL 外,官方还支持视频 Base64,以及 Files API 上传后的 stepfile:// 引用——适合不便公开直链的场景。详见模型页与 Video best practices。

四、读响应与计费提醒

非流式请求的读法与上一课相同:

response.choices[0].message.content

JSON 形状回顾:

{
  "model": "step-5-preview",
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": "The model's response"
      },
      "finish_reason": "stop"
    }
  ]
}

图片/视频会按 token 计入输入(与文本一样计费)。当前 step-5-preview 标价仍为:输入 cache-miss $1.00 / cache-hit $0.10 / 输出 $2.70(每 1M tokens)。多模态请求通常比纯短文本更费 token,建议先小图、短视频验证链路,再放大素材。

五、本课自检清单

  • model 仍是 step-5-preview
  • content 是数组,同时包含 type: text 与 type: image_url 或 type: video_url
  • 已替换所有 example.com 占位为你自己的可达 URL
  • 图片格式在 JPG/PNG/WebP/静态 GIF 内,单次 ≤ 60 张
  • 视频为 MP4/QuickTime/Matroska;MP4 URL 尽量 < 128 MB、< 5 分钟
  • 成功时仍从 choices[0].message.content 取正文

六、本课未覆盖(需要时再查官方)

  • 流式输出、Tool calling、JSON Mode / JSON Schema
  • reasoning_effort(low / medium / high)与 Messages API 的 output_config.effort
  • Prompt caching 细则、Claude Code / Step Plan 集成
  • 账号档位与 RPM/TPM 限流表(见 Pricing「Tiered Rate Limits」)

做完两课,你已经走通:密钥 → 文本 → 图/视频。上线前请再对照官方 API Reference 补超时、重试、错误处理与密钥管理。