WAND wiki
최근 문서

검색 결과가 없습니다. 모델명이나 다른 키워드로 검색해 보세요.

VOD APILast Updated 2026-09-05

PixVerse

On this page

Endpoint: POST https://vod.intl.tencentcloudapi.com
Action: CreateAigcVideoTask

기본 정보

항목 값
ModelName PixVerse
ModelVersion v6 / v5.6 / c1 (소문자)
기본값 ModelVersion=v6 / Resolution=1080P
가드레일 해제 지원 미지원

버전별 지원 규격

버전 해상도 비율 길이
공통 480P / 720P / 1080P / 2K / 4K — —
c1 360P / 540P / 720P / 1080P — 1–15초
v6 360P / 540P / 720P / 1080P — 1–15초
v5.6 360P / 540P / 720P / 1080P — 5 / 8 / 10초; 1080P는 10초 제외

입력 조건

버전 조건
c1 참조 / 오디오: 이미지 최대 7장, 영상 참조, 오디오 동시 생성
v6 참조 / 오디오: 이미지 최대 7장, 오디오 동시 생성
v5.6 참조 / 오디오: c1/v6의 참조 / 오디오 기능과 구분

요청 파라미터

파라미터 필수 타입 설명
ModelName 필수 String 고정값 PixVerse
ModelVersion 선택 String v6 / v5.6 / c1 (소문자)
Prompt 필수 String 생성 프롬프트
FileInfos.N 선택 Array 참조 입력. Usage는 FirstFrame(첫 프레임) 또는 Reference(참조)
OutputConfig.Resolution 선택 String 480P / 720P / 1080P / 2K / 4K
OutputConfig.Duration 선택 Integer 영상 길이(초)
OutputConfig.AspectRatio 선택 String 16:9 / 9:16 / 1:1 등

요청 예시

LANGUAGE
{
  "SubAppId": 123456789,
  "ModelName": "PixVerse",
  "ModelVersion": "v6",
  "Prompt": "a calm sunset over the ocean, cinematic",
  "OutputConfig": {
    "Resolution": "1080P",
    "Duration": 5,
    "AspectRatio": "16:9",
    "StorageMode": "Temporary"
  },
  "SessionContext": "job-001"
}

응답 예시

LANGUAGE
{
  "AigcVideoTask": {
    "Status": "FINISH",
    "ErrCode": 0,
    "Progress": 100,
    "Output": {
      "FileInfos": [
        {
          "FileUrl": "http://<host>.vod2.myqcloud.com/.../aigcVideoGenFile.mp4",
          "ExpireTime": "2026-08-01T10:29:48Z",
          "MetaData": {
            "Width": 1920,
            "Height": 1080,
            "Duration": 5.07,
            "Container": "mov,mp4,m4a",
            "Bitrate": 9494850
          }
        }
      ]
    }
  }
}

특수 설정

Lip Sync (립싱크)

인물이 말하는 소스 영상에 립싱크를 입힙니다. SceneType=lip_sync로 지정하고 음성을 두 가지 모드로 지정합니다.

TTS 모드: 대사와 스피커 지정

LANGUAGE
{
  "SubAppId": 123456789,
  "ModelName": "PixVerse",
  "ModelVersion": "lip_sync",
  "SceneType": "lip_sync",
  "Prompt": "talking",
  "FileInfos": [
    {
      "Type": "Url",
      "Category": "Video",
      "Url": "https://<cdn>/talking_head.mp4"
    }
  ],
  "ExtInfo": "{\"AdditionalParameters\":\"{\\\"lip_sync_tts_content\\\":\\\"안녕하세요, 반갑습니다.\\\",\\\"lip_sync_tts_speaker_id\\\":\\\"14\\\"}\"}",
  "OutputConfig": {
    "StorageMode": "Temporary"
  },
  "SessionContext": "job-001"
}

참조 오디오 모드: 오디오 파일 입력

LANGUAGE
{
  "SubAppId": 123456789,
  "ModelName": "PixVerse",
  "ModelVersion": "lip_sync",
  "SceneType": "lip_sync",
  "Prompt": "talking",
  "FileInfos": [
    {
      "Type": "Url",
      "Category": "Video",
      "Url": "https://<cdn>/talking_head.mp4"
    },
    {
      "Type": "Url",
      "Category": "Audio",
      "Url": "https://<cdn>/voice.mp3"
    }
  ],
  "OutputConfig": {
    "StorageMode": "Temporary"
  },
  "SessionContext": "job-001"
}

소스 영상은 공개 URL 사용

소스 영상 URL은 외부에서 접근 가능한 공개 URL이어야 합니다. 프리셋 스피커는 대부분 중국어 화자 기준이라, 한국어 대사는 참조 오디오 모드를 권장합니다.

버전 선택

ModelVersion 특징
c1 최신 버전. 캐릭터 일관성이 가장 좋고 참조 영상 입력과 립싱크를 지원
v6 범용 고품질 버전
v5.6 이전 세대 안정 버전. 영상 길이는 5 / 8 / 10초 중 선택

c1과 v6는 1~15초, v5.6은 5 / 8 / 10초만 지원합니다.

첫 프레임과 끝 프레임 지정

FileInfos.N.Usage로 첫 프레임과 끝 프레임을 구분합니다. LastFrame을 쓰는 방식을 권장합니다.

LANGUAGE
{
  "SubAppId": 123456789,
  "ModelName": "PixVerse",
  "ModelVersion": "v6",
  "Prompt": "Smooth transition",
  "FileInfos": [
    {
      "Type": "Url",
      "Category": "Image",
      "Url": "https://<cos>/first.jpg",
      "Usage": "FirstFrame"
    },
    {
      "Type": "Url",
      "Category": "Image",
      "Url": "https://<cos>/last.jpg",
      "Usage": "LastFrame"
    }
  ],
  "OutputConfig": {
    "StorageMode": "Temporary"
  },
  "SessionContext": "job-001"
}
Usage 의미
FirstFrame 첫 프레임
LastFrame 끝 프레임
Reference 참조 이미지

끝 프레임은 두 가지 방식이 있습니다

FileInfos.Usage=LastFrame을 쓰는 방식이 기준입니다. 최상위 LastFrameUrl도 하위 호환으로 동작하지만 새 구현에서는 Usage=LastFrame을 사용하세요. 끝 프레임 URL은 5MB 이하여야 합니다.

다중 이미지 참조 (Text / @Name)

참조 이미지마다 Text로 이름을 붙이고, 프롬프트에서 @Name으로 지목합니다. 이미지는 최대 7장입니다.

LANGUAGE
{
  "SubAppId": 123456789,
  "ModelName": "PixVerse",
  "ModelVersion": "c1",
  "FileInfos": [
    {
      "Type": "Url",
      "Category": "Image",
      "Url": "https://<cos>/character.png",
      "Text": "hero",
      "Usage": "Reference"
    },
    {
      "Type": "Url",
      "Category": "Image",
      "Url": "https://<cos>/fan.jpg",
      "Text": "fan",
      "Usage": "Reference"
    },
    {
      "Type": "Url",
      "Category": "Image",
      "Url": "https://<cos>/earring.jpg",
      "Text": "earring",
      "Usage": "Reference"
    }
  ],
  "Prompt": "the woman in @hero slowly raises her hand and opens the @fan, while the @earring sways gently as she turns her head",
  "OutputConfig": {
    "Duration": 8,
    "AspectRatio": "3:4",
    "AudioGeneration": "Enabled",
    "StorageMode": "Temporary"
  },
  "SessionContext": "job-001"
}

@Name 뒤에는 반드시 공백

@hero runs처럼 참조 대상 이름 뒤를 띄워 사용합니다. 프롬프트의 이름과 Text 값을 일치시키고 이름은 영문 / 숫자로 작성합니다.

참조 유형 (ReferenceType)

Category=Video인 항목에 ReferenceType을 지정해 참조의 성격을 정합니다. GV, Kling, PixVerse에 적용됩니다.

값 의미
subject 피사체 중심 참조
background 배경 중심 참조
LANGUAGE
{
  "SubAppId": 123456789,
  "ModelName": "PixVerse",
  "ModelVersion": "v5.6",
  "FileInfos": [
    {
      "Type": "Url",
      "Category": "Video",
      "Url": "https://<cos>/source.mp4",
      "ReferenceType": "subject"
    }
  ],
  "Prompt": "change the color of the main character's dress to white",
  "OutputConfig": {
    "StorageMode": "Permanent",
    "MediaName": "PixVerse-video-edit"
  },
  "SessionContext": "job-001"
}

립싱크 (lip_sync)

ModelVersion=lip_sync와 SceneType=lip_sync를 함께 지정합니다. 음성은 오디오 파일 또는 TTS로 지정합니다.

프리셋 스피커 목록

ExtInfo의 lip_sync_tts_speaker_id에 아래 값 중 하나를 지정합니다. 최신 목록은 PixVerse TTS 음색 조회 API로 확인할 수 있습니다.

speaker_id 설명
Auto 자동 선택
2 Zhen Youyu
4 외국인 남성
6 Li Jie
10 Jiang Jianghao
11 Lao Sen
12 Li Jieke
13 Qian Duoduo
14 Wang Xiaopai
16 시골 큰 목소리
18 허난 사투리
19 대만 억양
20 산시 억양
21 홍콩 억양
LANGUAGE
{
  "SubAppId": 123456789,
  "ModelName": "PixVerse",
  "ModelVersion": "lip_sync",
  "SceneType": "lip_sync",
  "FileInfos": [
    {
      "Type": "Url",
      "Category": "Video",
      "Url": "https://<cos>/talking_head.mp4"
    }
  ],
  "Prompt": "dancing",
  "ExtInfo": "{\"AdditionalParameters\": \"{\\\"lip_sync_tts_content\\\": \\\"lets dance and sing with me\\\", \\\"lip_sync_tts_speaker_id\\\": \\\"2\\\"}\"}",
  "OutputConfig": {
    "StorageMode": "Temporary"
  },
  "SessionContext": "job-001"
}

ExtInfo.AdditionalParameters 필드:

필드 설명
lip_sync_tts_content 입 모양에 맞출 대사
lip_sync_tts_speaker_id 프리셋 스피커 ID

입력 규격

항목 제약
이미지 크기 10MB 이하
이미지 포맷 jpeg / jpg / png
끝 프레임 URL 5MB 이하
참조 이미지 최대 7장

기타 파라미터

파라미터 설명
NegativePrompt 결과에서 제외할 요소를 설명하는 프롬프트
EnhancePrompt 프롬프트 자동 보정. Enabled / Disabled
Seed 난수 시드. 같은 값이면 결과를 재현할 수 있음
SessionContext 콜백에 그대로 전달되는 값. 최대 1000자
SessionId 중복 제거용 식별자. 3일 내 동일 ID 요청은 기존 요청 기준 처리