내레이션이 들어간 비디오 만들기
LLM으로 대본을 쓰고, 클립을 생성하고, 보이스오버를 합성해 하나로 합치는 엔드투엔드 과정 — API 키 하나로 끝냅니다.
이 튜토리얼은 하나의 스크립트 안에서 세 가지 모델 유형을 엮습니다. 언어 모델이 내레이션을 쓰고, 비디오 모델이 영상을 만들고, 음성 모델이 대본을 소리 내어 읽습니다. 하나의 API와 하나의 잔액으로 모델 유형을 넘나드는 것이 왜 유용한지 보여 주는 현실적인 예제입니다.
필요한 것: Python 3.9+, API 키, 그리고 마지막에 오디오와 비디오를 합치려면 ffmpeg.
이 튜토리얼은 실제 생성 작업을 실행하며 실제 크레딧을 소모합니다. 비디오가 가장 비싼 부분입니다 — 이것저것 시도하는 동안에는 짧은 길이로 시작하고, 정확한 금액을 미리 알고 싶다면 가격 추정을 이용하세요.
준비
pip install requests
export ATLASCLOUD_API_KEY="your-api-key"1단계 — 공통 헬퍼
모든 미디어 작업은 제출한 뒤 폴링하는 동일한 형태를 따르므로, 한 번만 정의해 둡니다.
import os, time, requests
API_KEY = os.environ["ATLASCLOUD_API_KEY"]
BASE = "https://api.atlascloud.ai/api/v1"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}
TERMINAL = {"completed", "succeeded", "failed", "timeout"}
def submit(endpoint: str, model: str, **params) -> str:
"""작업을 제출한다. 파라미터는 평평하게 두고 input으로 감싸지 않는다."""
r = requests.post(f"{BASE}/model/{endpoint}",
headers=HEADERS, json={"model": model, **params}, timeout=60)
r.raise_for_status()
return r.json()["data"]["id"]
def wait(prediction_id: str, timeout: int = 900) -> dict:
"""종료 상태까지 폴링하며 간격을 점점 늘린다."""
deadline, delay = time.time() + timeout, 2.0
while time.time() < deadline:
r = requests.get(f"{BASE}/model/prediction/{prediction_id}",
headers=HEADERS, timeout=30)
r.raise_for_status()
data = r.json()["data"]
if data.get("status") in TERMINAL:
if data["status"] in ("failed", "timeout"):
raise RuntimeError(f"작업 실패: {data.get('error') or data['status']}")
return data
time.sleep(delay)
delay = min(delay * 1.5, 10.0)
raise TimeoutError(prediction_id)
def download(url: str, path: str) -> str:
r = requests.get(url, timeout=300)
r.raise_for_status()
with open(path, "wb") as f:
f.write(r.content)
return path2단계 — LLM으로 대본 쓰기
언어 모델은 기본 URL이 다르고 동기 방식이므로 submit/wait을 거치지 않습니다.
def write_narration(topic: str) -> str:
r = requests.post(
"https://api.atlascloud.ai/v1/chat/completions",
headers={**HEADERS, "Content-Type": "application/json"},
json={
"model": "deepseek-ai/deepseek-v3.2",
"messages": [
{"role": "system",
"content": "You write narration for short videos. "
"Reply with two sentences of spoken narration and nothing else."},
{"role": "user", "content": f"Topic: {topic}"},
],
"max_tokens": 200,
},
timeout=120,
)
r.raise_for_status()
return r.json()["choices"][0]["message"]["content"].strip()
narration = write_narration("how ocean waves shape a coastline")
print(narration)3단계 — 영상 생성하기
video_id = submit(
"generateVideo",
"alibaba/wan-2.5/text-to-video",
prompt="slow aerial shot of waves breaking against a rocky coastline at golden hour",
duration=5,
)
video = wait(video_id)
download(video["outputs"][0], "footage.mp4")비디오 생성은 초가 아니라 분 단위로 걸립니다. 프로덕션에서는 웹훅을 등록해, 폴링 루프를 계속 열어 두는 대신 작업이 끝나면 콜백을 받으세요.
4단계 — 보이스오버 합성하기
2단계에서 만든 내레이션이 여기서 입력이 됩니다. 음성, 음악, 전사는 모두 같은 엔드포인트를 공유하며, 어느 것이 될지는 모델이 결정합니다.
audio_id = submit(
"generateAudio",
"bytedance/seed-audio-1.0",
text=narration,
format="mp3",
sample_rate=44100,
speech_rate=-5, # 조금 느린 편이 내레이션으로 알아듣기 좋다
)
audio = wait(audio_id)
download(audio["outputs"][0], "voiceover.mp3")음성 레퍼런스를 포함한 전체 파라미터는 오디오 모델을 참조하세요.
5단계 — 합치기
ffmpeg -i footage.mp4 -i voiceover.mp3 \
-c:v copy -c:a aac -shortest narrated.mp4선택 사항 — 자막 생성하기
보이스오버를 전사 모델에 다시 통과시키면 자막용 단어 단위 타이밍을 얻을 수 있습니다.
stt_id = submit(
"generateAudio",
"bytedance/seed-asr-2.0",
audio_url=audio["outputs"][0], # 필드 이름이 audio_url임에 주의
enable_punc=True,
show_utterances=True,
)
stt = wait(stt_id)
# 전사 모델에서 outputs[0]은 파일 URL이 아니라 텍스트 자체다
print(stt["outputs"][0])
for word in stt.get("stt_result", {}).get("words", [])[:10]:
print(word["start"], word["end"], word["text"])전사 모델은 audio_url을 받지만, 일부 다른 음성 인식 모델은 대신 audio를 받습니다. 해당 모델의 API 레퍼런스를 확인하세요 — 필드 이름을 잘못 넘기면 검증에서 실패합니다.
다음으로 해 볼 것
- 비디오 모델을 참조 이미지를 받는 모델로 바꾸고, 직접 제공한 스틸 이미지로 분위기를 결정해 보세요
- 폴링을 웹훅으로 옮겨 긴 작업이 프로세스를 막지 않게 하세요
- 여러 클립을 병렬로 실행한 뒤 이어 붙여 더 긴 시퀀스를 만드세요
- 오류와 요청 제한의 재시도 처리를 추가하세요
관련 문서
Last updated on