RZL

Blueprint terbuka

PABRIK

Pabrik Konten $1.5: satu kalimat topik jadi video esai 10 menit, tanpa tangan manusia kecuali satu klik.

v1.0 — 22 Juli 2026

Apa yang jalan, dan apa buktinya

Halaman ini menjelaskan satu pipeline video yang jalan di server saya sendiri: dari satu kalimat topik jadi esai video ~10 menit lengkap dengan narasi, klip film, motion graphic, subtitle, judul, deskripsi, dan thumbnail. Bukti yang bisa saya tunjuk cuma satu, dan saya tulis apa adanya: render contoh terakhir yang tercatat di dokumentasi internal saya berdurasi 585 detik dengan skor QC teknis 100/100, dan biaya per video di kisaran $1.5–2 — mayoritas dari itu biaya suara.

Setiap angka di halaman ini punya penanda. Yang terukur dari render nyata ditandai terukur; yang cuma perkiraan struktural ditandai perkiraan dan tetap saya biarkan jadi perkiraan, bukan dibulatkan jadi fakta. Dari empat mesin yang jalan di server ini, hanya satu yang biayanya terukur. Tiga sisanya perkiraan, dan kalau kamu menyusun anggaran dari angka itu, kamu menyusun anggaran dari tebakan saya.

Prompt produksi di bawah disalin apa adanya dari sistem yang jalan, bukan versi yang dirapikan untuk halaman ini — termasuk kalimat yang berantakan dan aturan yang kelihatan berlebihan, karena justru aturan-aturan itulah hasil dari kegagalan sebelumnya. Tidak ada satu pun fragmen yang saya ganti placeholder: setelah dipindai, prompt-prompt ini memang tidak memuat host internal, alamat IP, path server, nama container, atau kredensial. Yang rahasia ada di environment variable dan credential store n8n, bukan di teks prompt. Satu pengecualian bentuk, bukan isi: baris pemisah seperti '### PASS 1 ... ###' dan label tiap mood di prompt musik saya tambahkan sendiri supaya beberapa prompt yang digabung di satu blok tetap terbaca, dan satu nilai (kecepatan bicara narasi, angka desimal) saya isi ke slot yang di kode sumbernya berupa f-string — teks instruksinya sendiri tidak saya ubah satu kata pun.

Ditulis untuk orang yang sudah pernah menjalankan n8n, punya VPS atau server sendiri, dan terbiasa menaruh API key di environment. Bukan untuk orang yang mencari layanan siap pakai atau tombol satu klik. Tidak ada apa pun di halaman ini yang jalan tanpa server yang kamu urus sendiri, dan sebagian besar waktu yang saya habiskan bukan untuk membangunnya, melainkan untuk mencari tahu kenapa sesuatu diam-diam tidak jalan.

Empat mesin

MesinJenisDurasiTriggerBiaya/video
WF6 CinematicEsai film sinematik, gaya dokumenter~10 menitOn-demand — Telegram, atau POST /webhook/cinematic-trigger~$1.5–2 per video
WF7 ToonAnimasi kartun edukasi finansial (Remotion)~8+ menit, plus versi pendek vertikalTelegram: toon <topik>~$1–2 per video (perkiraan)
StoryVideo cerita pendek vertikal~60–79 detikOtomatis dari daftar topik, atau on-demand~$0.30–0.60 per video (perkiraan)
ClippingPotongan viral dari video panjangShorts vertikalTelegram, atau riset otomatis~$0.20–0.50 per klip (perkiraan)

Bongkar biaya

KomponenBiayaCatatan
Suara narasi (ElevenLabs, suara kloning)~$1.5Penyumbang terbesar — mayoritas dari total. Satu naskah ~1.500–1.900 kata dibacakan penuh.
Keputusan yang memotongnya separuh: model VO~$3 → ~$1.5Pindah dari eleven_multilingual_v2 ke eleven_flash_v2_5. Satu baris environment variable (CLIP_ELEVEN_MODEL), tanpa rebuild image, kualitas tetap dipakai produksi.
Penulisan naskah (2 panggilan Claude Sonnet 5)beberapa puluh senDraft plain-text lalu refine jadi JSON. Angka ini sisa dari total nyata dikurangi VO — bukan tagihan yang dibaca per baris.
Klip film (Clip Cafe)$0 marginalMasuk kuota langganan ~900 klip/bulan. Gratis sampai kuota habis, lalu fail-closed: workflow berhenti, bukan diam-diam merender tanpa klip.
Musik latar — jalur pool CC-BY$0Lagu berlisensi CC-BY dengan atribusi otomatis di deskripsi. Ini jalur yang dipakai saat angka ~$1.5–2 dicatat.
Musik latar — jalur generated (default sejak 18 Jul 2026)~$2.16 per render untuk ~6 render pertama, lalu $0 (perkiraan)Perkiraan, belum terukur dari render nyata — dihitung dari batas yang dipaksakan kodenya, bukan dari log render. bgm_bed.py menjalankan maksimal satu top-up per babak, tiap top-up GEN_SEGMENT_SECONDS=45 detik, dan melewati top-up yang tidak muat penuh di sisa jatah — dengan GEN_BGM_MAX_NEW_SEC_PER_RENDER=300, jatah itu cuma memuat 6×45=270 detik penuh, bukan 300. Di tarif elevenlabs/music $0.48/menit (cinematic_producer.py), 270 detik ≈ $2.16. Biayanya TIDAK turun mulai render kedua: gen_bgm_min_pool=6 berlaku PER MOOD, dan arc 7 babak memetakan ke 6 mood berbeda (GENERATED_MOOD_MAP) di setiap render, jadi kolamnya baru penuh (6 mood × 6 track = 36 track) setelah kira-kira 6 render berturut-turut @ ~$2.16 (~$13 total) — baru sesudah itu jatuh ke $0. Kembali gratis kapan saja dengan CINE_BGM_MODE=pool.
Render, transkrip, encode (FFmpeg + Whisper)$0 marginalSemua di server sendiri. Tidak ada biaya cloud per video; listrik tidak dihitung.
Total per video ~10 menit~$1.5–2Angka jalur pool, bukan jalur musik generated. Render contoh terakhir yang tercatat: 585 detik, skor QC teknis 100/100.

Tahap demi tahap

1. Kirim satu kalimat topik

Menentukan subjek esai — judul film, judul serial, atau topik finansial bebas.

Tool: n8n webhook (POST /webhook/cinematic-trigger) atau bot Telegram · Biaya: $0

Jebakan: Field-nya bernama movie_title demi kompatibilitas router lama walau isinya boleh topik bebas, dan kalau body-nya tidak memuat field itu sama sekali selector diam-diam jatuh ke seed harian (indeks = day-of-year) — request salah bentuk tetap menghasilkan video, cuma bukan topik yang kamu minta.

2. Tulis naskah dua kali: draft, lalu perbaiki

Menghasilkan esai 7 babak dengan alur emosi, plus metadata terstruktur (judul, tag, kutipan kinetik, beat motion graphic) dalam satu JSON.

Tool: Claude Sonnet 5, dua panggilan HTTP ke API Anthropic — max_tokens 16384 (draft) dan 20000 (refine), thinking:{type:'disabled'} di keduanya · Biaya: beberapa puluh sen untuk dua panggilan

### PASS 1 — system prompt panggilan draft (output teks biasa) ###

You are a world-class FINANCE video-essay writer. Write a ROUGH DRAFT for a long-form YouTube video essay on the given SUBJECT. Plain text, NOT JSON.

LANGUAGE: Compose the narration prose — WORKING TITLES, OUTLINE lines, and especially the FULL DRAFT NARRATION — natively in cinematic BAHASA INDONESIA: natural, mature, spoken-word register, NOT translationese and not stiff formal EYD. Do not draft in English and lean on the next pass to translate it; write it AS Indonesian prose from the first word. QUOTE CANDIDATES entries — whether iconic film dialogue or your own sharpest original lines — must be written in ENGLISH; only the FULL DRAFT NARRATION, WORKING TITLES, and OUTLINE are Bahasa Indonesia.
Keep the STRUCTURAL SCAFFOLDING below in English exactly as printed — it is relied on downstream and must not be translated or reworded: the section headers themselves (SUBJECT KIND, WORKING TITLES, OUTLINE, FULL DRAFT NARRATION, QUOTE CANDIDATES, THUMBNAIL IDEAS), the 7-act arc labels (HOOK / WORLD / DESCENT / BREAKING POINT / REVELATION / TRANSFORMATION / MIRROR), and above all the SUBJECT KIND answer itself — it MUST be the literal English word 'movie', 'series', or 'topic', never translated to 'film', 'serial', or 'topik'.

Include, in this order:
1. SUBJECT KIND: movie | series | topic (one word).
2. WORKING TITLES: 5 candidate YouTube titles (<=60 chars each, curiosity gap + subject name).
3. OUTLINE: the 7-act arc (HOOK / WORLD / DESCENT / BREAKING POINT / REVELATION / TRANSFORMATION / MIRROR), one line each.
4. FULL DRAFT NARRATION: every act written out. HARD word floor by kind: topic >=1500, movie >=1800, series >=3000 words total. Go long and raw — the refine pass edits you down, it cannot invent depth.
5. QUOTE CANDIDATES: 10-14 short quotable money lines (<=14 words) with attribution — iconic film dialogue or your sharpest own lines.
6. THUMBNAIL IDEAS: 3 options — text (<=5 words) + which visual moment of the essay it pairs with.
Never plot-recap voice, never generic financial advice.

### PASS 2 — system prompt panggilan refine (output JSON final) ###

You are a world-class FINANCE video-essay writer — in the style of the best money-psychology essays: moody, cinematic, insightful, never a plot recap and never generic financial advice.

LANGUAGE: Write the ENTIRE narration in cinematic BAHASA INDONESIA — natural, mature, spoken-word register; not translationese, not formal EYD stiffness. Indonesian finance vocabulary where it exists (imbal hasil, arus kas, bunga berbunga); keep the English term only where it is what people actually say (cash flow, compounding, leverage).
EXCEPTIONS that stay in ORIGINAL ENGLISH — do NOT translate:
  • every 'quote' field (film dialogue on kinetic cards)
  • 'movie_title', 'clip_films', and any proper noun
  • 'scene_phrases' (they are matched against English film subtitles)
The 'title' and 'description' fields are Bahasa Indonesia (Indonesian YouTube audience); hashtags may stay English.

NARASI = BAHASA LISAN untuk DIUCAPKAN, bukan artikel dibacakan: kalimat pendek (rata-rata di bawah 14 kata), sapaan langsung ke penonton, pertanyaan retoris, ritme variatif — kalimat 3 kata boleh disusul kalimat 18 kata. Tes: baca keras; kalau terdengar seperti membaca teks, tulis ulang.
VO PERFORMANCE TAGS (mesin TTS mengerti ini, taruh DI DALAM narration): [pause] untuk jeda dramatis (paling kuat tepat sebelum punchline atau angka mengejutkan), <soft>...</soft> untuk kalimat intim/pelan, <loud>...</loud> untuk penekanan, JARANG [laugh]. Aturan keras: tag pembungkus harus buka+tutup di DALAM SATU kalimat; maksimal 1 tag per 2 kalimat; TANPA tag di title/description/kinetic_quotes/gfx_beats (itu teks tampil, bukan VO).

You are given a SUBJECT. Decide the ESSAY MODE and emit it as 'essay_mode':
- If the SUBJECT is a specific film/series title: essay_mode='film'. The essay analyzes THAT work; every act's movie_title is that exact title and clip_films MUST be an empty array — clips never stray to other films.
- If the SUBJECT is a finance TOPIC (e.g. 'why the rich get richer and the poor stay poor'): essay_mode='topic'. The narration is topic-driven; illustrate each act with scenes from WHATEVER films fit best — any film, any genre (Parasite, Nightcrawler, The Pursuit of Happyness...), not just classic finance movies. Set each act's movie_title to that act's MAIN film and add up to 3 clip_films per act. Set top-level movie_title to the most central film of the essay.

Also emit 'subject_kind': 'movie' | 'series' | 'topic'. A TV series (Ozark, Succession, Billions...) is subject_kind='series' with essay_mode='film' (clips still lock to that one series).

You are given your own ROUGH DRAFT from a previous pass. Silently critique it — pacing, hooks, open loops, specificity, word floor, retention — then REWRITE it into the polished FINAL script: keep the sharpest lines, cut flab, upgrade weak transitions. Output the final JSON only, never the critique.

LENGTH — HARD MINIMUM word floors at ~150 wpm (longer is fine, shorter is a failure):
- subject_kind='topic': >=1500 words (>=10 min), target 1500-1900.
- subject_kind='movie': >=1800 words (>=12 min), target 1800-2200.
- subject_kind='series': >=3000 words (>=20 min), target 3000-3600 — a series gives you seasons of material; go deep, arc by arc.

Write a LONG-FORM essay about MONEY, using that FILM/SERIES as the vehicle. The essay must extract a FINANCIAL and PSYCHOLOGICAL lesson the viewer can apply — greed, risk, leverage, incentives, status games, wealth vs. income, the price of ambition. The film is the case study; the viewer's money life is the subject. Assume the viewer has seen trailers but maybe not the film; spoil only what the lesson requires and SAY so casually when you do.

7-ACT STRUCTURE (write every act, in order):
ACT1 THE HOOK (0:00-0:40, ~180 words): open on the film's most arresting money moment or line, then frame the REAL financial question of the essay (about the viewer's life, not the plot). Open loop. NEVER 'Did you know'.
ACT2 THE WORLD (~280 words): the character and their relationship with money — what they want, what they tell themselves it will buy, why we recognize our own financial self-deception in them.
ACT3 THE DESCENT (~280 words): the central financial conflict / slippery slope — specific deals, trades, debts, choices, turning points.
ACT4 THE BREAKING POINT (~280 words): the darkest beat — the margin call, the bust, the moral bill arriving. Make the viewer feel it.
ACT5 THE REVELATION (~220 words): the film's core financial insight — name the real concept precisely (moral hazard, sunk cost, leverage, hedonic treadmill, principal-agent problem, asymmetric risk...).
ACT6 THE TRANSFORMATION (~250 words): how the character changes or refuses to — and what it costs them in money AND meaning. Contrast with the alternative path.
ACT7 THE MIRROR (~180 words): turn it on the viewer — the applicable money lesson for their life, landing cinematically, then a warm VARIED subscribe CTA woven into the reflection (also emit in 'outro_cta'). Keep act_title exactly 'THE MIRROR'.

The per-act word counts above are the ~1670-word TOPIC baseline — SCALE every act proportionally to hit your subject_kind's word floor (movie ~1.15x, series ~1.9x). Always exactly 7 acts.

EMOTIONAL ARC: curiosity -> admiration -> empathy -> pain -> hope -> inspiration -> reflection.
RETENTION CRAFT: specific scenes with episode/season or act references; 3-5 open loops; pattern interrupt every ~100 words; 2-3 rhetorical questions per act; occasional second person; each act ends with a cliffhanger into the next. NO plot-summary voice, NO get-rich-quick tone, NO motivational cliches.

OUTPUT: VALID JSON ONLY, EXACTLY this shape:
{
  "title": "YouTube title, max 60 chars, film name + money curiosity gap",
  "description": "hook line + value promise + chapter timestamps + 3-5 hashtags",
  "tags": ["tag1","tag2"],
  "essay_mode": "film",
  "subject_kind": "movie",
  "movie_title": "the exact film/series title",
  "movie_year": 1987,
  "script_clean": "full script text, all 7 acts joined in order",
  "word_count": 1950,
  "estimated_duration_sec": 780,
  "emotional_arc": ["curiosity","admiration","empathy","pain","hope","inspiration","reflection"],
  "bgm_mood": "cinematic_dark_crescendo",
  "thumbnail_text": "max 5 words",
  "thumbnail_moment": 0.62,
  "outro_cta": "warm varied subscribe line",
  "acts": [
    {"act_number":1,"act_title":"THE HOOK","text":"...","movie_title":"the exact film/series title","caption_query":"semantic description of the scene wanted, e.g. Gekko pitches greed to the shareholders","scene_phrases":["exact short dialogue quote","another iconic money line"],"kinetic_quotes":[{"quote":"short punchy money quote, max 14 words","attribution":"Gordon Gekko"}],"gfx_beats":[{"type":"counter","value":2300000000,"unit":"$","prefix":"-","caption":"kerugian nasabah 3 bulan"},{"type":"line_chart","title":"Harga saham","series":[{"label":"1998","value":12},{"label":"2001","value":0.26}],"unit":"$","caption":"dari 12 dolar ke 26 sen"},{"type":"icon_metaphor","icon":"scale","caption":"regulator vs bank"},{"type":"node_flow","layout":"horizontal","nodes":["Nasabah","Broker","Offshore"],"caption":"ke mana uang mengalir"},{"type":"headline","text":"Semua orang tahu. Tidak ada yang bicara.","emphasis":["tahu"]},{"type":"doc_card","headline":"Bank Kolaps Semalam","body":"Ribuan nasabah antre di depan kantor cabang.","emphasis":["Semalam"]}],"clip_films":[{"movie_title":"Margin Call","movie_year":2011,"caption_query":"the firm decides to dump the toxic assets","scene_phrases":["be first be smarter or cheat"]}],"broll_queries":["metaphorical stock query","another"],"diegetic_ok":true,"emotional_tone":"curiosity","music_intensity":"low","subtitle_style":"main"}
  ],
  "youtube_chapters": [ {"time":"0:00","title":"The Hook"} ]
}

FIELD RULES:
- acts: EXACTLY 7, all with the keys shown above.
- title: pick the BEST of your draft's candidate titles, sharpened. Curiosity gap + the film/series/topic name, <=60 chars, specific stakes ('Wall Street Explained' = bad; 'Gordon Gekko Was Right About One Thing' = good). Never generic, never a lie the video can't cash.
- description: line 1 = the hook question ALONE (it shows above the fold), then 2-3 sentences of what the viewer will learn, then the chapter timestamps, then 3-5 niche hashtags.
- thumbnail_text: <=5 words, HIGH tension, complements (never repeats) the title — the two together form one curiosity gap.
- thumbnail_moment: fraction 0.0-1.0 of the video runtime where the most ARRESTING visual for a thumbnail frame lands (usually the ACT4 breaking point ~0.55-0.70, or the hook scene ~0.03-0.08). The renderer grabs the thumbnail frame there.
- word_count: your actual count — MUST meet the subject_kind floor.
- movie_title: EXACT title on every act (used to fetch real clips of THAT film — never a different film).
- caption_query: ONE plain-English semantic description of the scene(s) wanted for this act — prefer scenes about money, deals, wealth, ruin.
- scene_phrases: 3-4 SHORT VERBATIM dialogue quotes from THIS film/series (<=8 words each) matching the act's beat — PREFER lines about money/greed/risk; they are transcript search keys for real clips, accuracy matters more than beauty.
- kinetic_quotes: 1-2 per act. SHORT (<=14 words) quotable money lines — the film's iconic financial dialogue, or a sharpened line from YOUR essay. attribution = speaker or film title. These render as typography cards between clips. ACT1's first kinetic_quote is the HOOK card (make it the sharpest line of the essay); ACT7's first is the CTA/mirror card.
- gfx_beats: 3-4 per act, WAJIB. Ini motion graphics gaya dokumenter Vox yang mengisi layar di antara movie clips. type: counter (SATU angka dramatis), line_chart (2-8 titik data), icon_metaphor (icon: wallet|coins|banknote|piggy-bank|trending-up|trending-down|chart-line|house|car|smartphone|coffee|shopping-cart|alarm-clock|shield|umbrella|scale|hourglass|rocket|flame|gem), node_flow (alur uang/sebab-akibat/timeline, 2-5 node), headline (kalimat paling tajam act itu + emphasis 1-2 kata), doc_card (potongan 'koran' — headline + 1-2 kalimat body). SEMUA angka HARUS berasal dari narasinya sendiri — angka karangan = gagal. caption max 10 kata, Bahasa Indonesia.
gfx_beats WAJIB UNIK antar act: jangan ulangi angka, judul headline, icon, chart, atau node_flow yang sudah dipakai act lain — setiap beat harus data atau metafora BARU. kinetic_quotes juga unik antar act (jangan mengulang quote yang sama).
- essay_mode: 'film' or 'topic' (see MODE rules above). This gates clip scoping in the renderer — be precise.
- clip_films: OTHER films whose real clips fit this act's beat. essay_mode='film' → MUST be []. essay_mode='topic' → up to 3 per act, any film that serves the narration; scene_phrases are verbatim quotes from THAT film.
- diegetic_ok: true for acts where real film clips should dominate (most acts); false only if an act is pure abstract reflection.
- broll_queries: metaphorical stock queries (bad: 'man sad', good: 'empty trading floor at night') — LEGACY fallback only, still required.
- emotional_tone one of: curiosity, admiration, empathy, pain, hope, inspiration, reflection. subtitle_style one of: emphasis, main, conflict. music_intensity one of: low, medium, high.
- script_clean MUST equal all 7 act texts concatenated in order.
- youtube_chapters: one per act, cumulative 'M:SS' at ~150 wpm.
- bgm_mood one of: cinematic_dark_crescendo, cinematic_melancholic, cinematic_inspiring, cinematic_tense.

Jebakan: Dua-duanya wajib mematikan thinking secara eksplisit dan memasang max_tokens besar; kalau tidak, JSON pass-2 terpotong di tengah dan node parser di belakangnya cuma bisa menyelamatkan output yang kelebihan teks, bukan yang kurang.

3. Bikin suara narasi

Membacakan naskah dengan suara kloning, sehingga terdengar seperti satu orang yang konsisten di semua video.

Tool: ElevenLabs, model eleven_flash_v2_5 — engine dipilih lewat CINE_TTS_ENGINE (env khusus modul cinematic; CLIP_ELEVEN_MODEL di baris BIAYA di atas punya prefix CLIP_ karena dipakai bareng modul clip_worker lain — dua prefix, satu sistem TTS, bukan dua sistem berbeda), bukan lewat kode. Panggilannya sendiri sekarang lewat proxy elevenlabs/tts di inference.sh, bukan langsung ke api.elevenlabs.io — tarif resmi ElevenLabs tidak langsung berlaku di sini · Biaya: ~$1.5 — porsi terbesar dari total

Jebakan: Tag performa VO ([pause], <soft>...</soft>, <loud>...</loud>) hanya boleh berada di dalam teks narasi dan harus buka-tutup dalam satu kalimat; kalau bocor ke judul, deskripsi, atau kartu kutipan, tag mentahnya ikut terbaca atau ikut tampil di layar.

4. Rencanakan shot, lalu ambil klip filmnya

Memecah naskah jadi rencana beat per beat (klip film / motion graphic / footage nyata / soundbite), lalu mengunduh klip yang cocok untuk tiap beat.

Tool: Claude Sonnet 5 sebagai shot planner (maksimal 2 panggilan per job, dibatasi keras di kode) + validator Python + footage miner ke Clip Cafe · Biaya: $0 marginal untuk klip (kuota langganan). Panggilan shot planner tidak dipisah dalam angka biaya dokumentasi — perlakukan sebagai tambahan kecil di atas biaya naskah.

You are a documentary-grade shot planner for a finished Bahasa Indonesia
video-essay narration script. You do NOT write narration — it already exists. You break it
into a beat-by-beat shot plan for the footage miner and motion-graphics renderer.

COMPOSITION BUDGET (percent of total runtime, self-report your totals so you can check them
before answering):
  - MOVIE_CLIP (finance/Wall-Street film clips): 38-45%
  - MGFX (motion graphics / data cards / kinetic text): 25-32%
  - REAL_FOOTAGE + SOUNDBITE combined (real archival/stock footage, SOUNDBITE counts toward
    this bucket since it IS real footage, just a quote-driven clip): 25-32%

SHOT LENGTH: every beat's est_duration_s <= 7s, EXCEPT shot_type MGFX which may run up to 10s.
If a narration sentence's natural duration would exceed the cap, SPLIT it into consecutive
beats at a sentence boundary rather than emitting one over-length beat.

BEAT COUNT IS A HARD OUTPUT-BUDGET CONSTRAINT — see the beat-count target in the user message
(it's specific to this script's runtime): your response has a fixed token ceiling, and each
beat's fixed JSON fields (shot_type, search_keywords, mgfx_template, mood, tier_max, etc.) cost
far more tokens than the narration text itself — so MINIMIZE the number of beats by running
each one as close to its max duration as the visuals allow. Group multiple consecutive
sentences that share the same shot/visual into ONE beat near the 7s (or MGFX's 10s) cap rather
than one beat per sentence. Only cut a beat short when the shot genuinely must change (new
visual, new must_show_number, a soundbite). A beat plan with one beat per short sentence is
WRONG even if each beat is individually valid — it will not fit the output budget.

SOUNDBITE beats (2-4 total, no more no less): placed at emotional pivots — end of babak 2,
babak 5 ("tamparan realita" / reality-slap moment), and pre-conclusion (late babak 6 or 7).
Each SOUNDBITE beat MUST set:
  - soundbite_quote_intent: prose description (English is fine) of the ideal quote for the
    footage miner to search for — what line, what emotional register, what kind of scene.
  - subtitle_id: the Bahasa Indonesia subtitle GLOSS (translation text) of the quote you expect
    the miner to find — this is what gets burned in as the Indonesian subtitle under the
    original-language soundbite audio. Never leave this null for a SOUNDBITE beat.
For every OTHER shot_type, subtitle_id MUST be null (only SOUNDBITE beats carry a gloss).

COLD OPEN: the very FIRST beat in the array MUST be beat_id "b00", babak 0, appearing before
any babak 1 beat. It carries NO narration (vo_text must be empty) — it is the strongest
soundbite or movie moment, a hook before the essay's first spoken word. shot_type must be
SOUNDBITE or MOVIE_CLIP.

PACING: narration duration is NOT estimated from a fixed words-per-second guess — this
production's Indonesian VO measures 2.2383 words/sec (measured from the actual
narration audio). Derive each non-cold-open beat's est_duration_s from its vo_text word count
at that rate, then adjust for splits.

BEAT COVERAGE: every word of every act's narration text must appear in exactly one beat's
vo_text, in order, babak matching act_number. Do not paraphrase, skip, or add narration text
that isn't in the source script.

ShotBeat schema (every beat, exactly these keys):
{
  "beat_id": "b23", "babak": 4, "vo_text": "...", "est_duration_s": 6.5,
  "shot_type": "REAL_FOOTAGE | MOVIE_CLIP | MGFX | SOUNDBITE",
  "search_keywords": ["bernanke hearing 2008", "congress testimony"],
  "mgfx_template": null, "mood": "tension", "must_show_number": "700 miliar dolar",
  "soundbite_quote_intent": null, "tier_max": 1, "subtitle_id": null
}
(tier_max: highest copyright tier this beat may use, 1 or 2 — Tier 3 is a miner-side
escalation under its own caps, never planner-assigned.)

OUTPUT COMPACTNESS (every character here is billed against your token ceiling):
  - search_keywords: EXACTLY 2 items, each <=4 words. Never 3+.
  - No pretty-printing: no blank lines or extra whitespace between beat objects, one compact
    JSON blob.

Return ONLY valid JSON, no markdown fences, no prose: {"beats": [...], "allocation": {...}}

Jebakan: Budget komposisi (38–45% klip film, 25–32% motion graphic, 25–32% footage nyata) dan batas panjang shot dipaksakan oleh validator Python, bukan dipercayakan ke model — rencana 'satu beat per kalimat' itu valid secara skema tapi tidak akan muat di plafon token, dan hasilnya adalah plan yang terpotong.

5. Render motion graphic

Mengisi layar di antara klip film dengan angka besar yang menghitung, grafik garis, ikon metafora, alur node, headline, dan kartu koran.

Tool: Renderer motion graphic sendiri di server (FFmpeg + subtitle/ASS), sumber datanya field gfx_beats dari JSON naskah · Biaya: $0 marginal — dirender di server sendiri

Jebakan: Semua angka di layar harus berasal dari narasinya sendiri dan tidak boleh mengulang antar babak; angka yang dikarang model lolos render tanpa error satu pun dan baru ketahuan salah oleh penonton yang mengeceknya.

6. Jahit semua jadi satu video

Menggabungkan klip, motion graphic, narasi, musik, dan subtitle jadi satu file 1920×1080 siap unggah.

Tool: FFmpeg untuk potong/encode, Whisper (lokal, bahasa dipin ke id) untuk subtitle kata-per-kata, timeline_assembler untuk audio bed dan musik · Biaya: $0 untuk render dan transkrip. Musik: $0 di jalur pool CC-BY; jalur generated ~$2.16/render untuk ~6 render pertama (lihat rincian di baris BIAYA musik generated), baru $0 setelah kolam per mood penuh.

### Katalog mood — satu prompt per mood, dipilih dari emotional_tone act ###

cinematic_dark_crescendo:
"A driving, modern documentary instrumental that builds tension with a rhythmic pulse — deep strings, ticking percussion, energetic momentum, dark but never sluggish. No vocals."

cinematic_melancholic:
"A reflective documentary instrumental with gentle forward motion — soft piano over a subtle rhythmic pulse, wistful but warm, never a dirge. No vocals."

cinematic_inspiring:
"An upbeat, optimistic documentary instrumental — bright piano, warm strings, driving rhythmic pulse, confident forward energy. No vocals."

cinematic_inspiring_build:
"An upbeat documentary instrumental building from a light rhythmic groove to a triumphant, energetic swell — strings, piano, punchy percussion. No vocals."

cinematic_tense:
"A tense but energetic documentary instrumental — pulsing low strings, urgent ticking percussion, driving momentum, suspense with drive. No vocals."

cinematic_mysterious_tense:
"A mysterious documentary instrumental with a steady rhythmic undercurrent — sparse textures over a driving pulse, intrigue with energy. No vocals."

### Dipakai kalau mood tidak dikenal ###

"An energetic, modern documentary instrumental with rhythmic pulse and forward drive. Strings, piano, percussion. No vocals."

Jebakan: Ada lebih dari satu modul yang terlihat mengurus musik dan sebagian besar sudah mati; jalur yang benar-benar jalan adalah timeline_assembler — grep log render dulu sebelum mengubah apa pun di tahap ini, karena kode yang kelihatan otoritatif di sini pernah ternyata dead code.

7. Susun judul, deskripsi, tag, dan thumbnail

Merakit paket SEO YouTube: judul terpotong aman, deskripsi dengan chapter, tag, hashtag, dan frame thumbnail.

Tool: Python murni (seo_generator) — tidak ada panggilan AI di tahap ini; bahannya dari JSON naskah tahap 2 · Biaya: $0 — tanpa panggilan model

Jebakan: Timestamp chapter yang ditebak di hulu sering meleset, jadi dipakai hanya kalau lolos aturan YouTube (chapter pertama di 0:00, jarak minimal 10 detik, minimal 3 chapter); kalau gagal, waktunya dihitung ulang proporsional dari jumlah kata tiap babak terhadap durasi audio yang sebenarnya.

8. Minta persetujuan sebelum publish

Mengirim preview ke Telegram dengan dua tombol; tidak ada yang terunggah tanpa satu klik manusia.

Tool: n8n Telegram node dengan tombol URL, memanggil webhook approve/reject (?run_id=&action=) · Biaya: $0

Jebakan: Gerbangnya sengaja menerima dua status, done dan needs_review, supaya render yang skornya jelek tetap sampai ke mata manusia dan bukan dibuang otomatis; yang benar-benar dihentikan hanya kegagalan keras — status lain, atau durasi di bawah 20 detik.

Tembok gotcha

Enam kegagalan nyata. Masing-masing memakan berhari-hari sebelum ketemu sebabnya.

n8n 2.x memblokir $env di dalam node

Gejala: Ekspresi yang membaca environment variable mengembalikan "access to env vars denied". Workflow lain di server yang sama jalan normal, jadi kelihatannya seperti masalah workflow itu sendiri.

Sebab: n8n 2.x memasang N8N_BLOCK_ENV_ACCESS_IN_NODE dengan default true. Hanya workflow yang benar-benar memakai $env yang terkena, jadi masalahnya tampak acak.

Solusi: Set N8N_BLOCK_ENV_ACCESS_IN_NODE=false di docker-compose.yml, lalu `docker compose up -d n8n`. Perhatikan: file compose ini biasanya BUKAN bagian dari repo aplikasi — mengubah repo tidak akan mengubah apa pun.

IF node bertipe number me-route ke FALSE diam-diam

Gejala: Cabang TRUE tidak pernah jalan. Tidak ada error, tidak ada log, eksekusi tercatat sukses.

Sebab: Perbandingan bertipe number pada nilai yang sampai sebagai string tidak cocok, dan n8n memperlakukan ketidakcocokan tipe sebagai FALSE, bukan sebagai error.

Solusi: Samakan tipe di node sebelumnya secara eksplisit, dan uji kedua cabang dengan data nyata — bukan dengan data contoh yang tipenya kebetulan sudah benar.

Dockerfile meng-COPY module satu per satu

Gejala: Fitur baru sama sekali tidak berpengaruh di produksi. Kode ada di repo, image ter-build tanpa error, container jalan.

Sebab: Dockerfile menyebut file sumber satu per satu. Module baru tidak ikut tersalin, dan import-nya gagal di dalam blok try/except yang menelan error — jadi fiturnya jadi no-op senyap, bukan crash.

Solusi: Setiap kali menambah module, periksa Dockerfile. Jangan biarkan ImportError ditelan diam-diam; minimal log peringatan supaya kegagalan terlihat.

versionId / activeVersionId desync

Gejala: Workflow menjalankan versi lama. Editor menampilkan perubahanmu, eksekusi memakai yang lama, dan tidak ada error apa pun.

Sebab: Mengedit workflow langsung lewat database tanpa menyinkronkan penunjuk versi aktif. n8n memuat versi yang ditunjuk penunjuk itu, bukan yang terbaru.

Solusi: Kalau menulis ke DB secara langsung, perbarui juga penunjuk versi aktif, lalu restart container n8n. Lebih aman: pakai REST API dan jangan sentuh DB sama sekali.

Auto-thinking menembus max_tokens

Gejala: JSON dari model terpotong di tengah. Kadang berhasil, kadang tidak — tergantung panjang jawaban.

Sebab: Mode thinking yang aktif otomatis memakan anggaran token yang sama dengan output. Anggaran habis sebelum JSON-nya selesai ditulis.

Solusi: Matikan thinking secara eksplisit untuk panggilan yang menuntut output terstruktur, dan minta output lewat tool-use, bukan lewat teks bebas yang harus di-parse.

Filter starvation — pipeline sehat, kolamnya yang kosong

Gejala: Pipeline berhenti menghasilkan apa pun. Semua node hijau, tidak ada error, jumlah hasil nol.

Sebab: Filter di hulu (durasi, lisensi, jumlah view, kesegaran) menyaring habis kandidat. Kolam sumbernya memang tidak cukup besar untuk memenuhi semua kriteria sekaligus.

Solusi: Ukur kolamnya sebelum menyalahkan kodenya: query sumber dengan tiap filter dinyalakan satu per satu, dan lihat di filter mana angkanya jatuh ke nol. Lalu pasang alarm yang berbunyi saat hasil nol — hasil nol harus berisik, bukan senyap.

Kenapa masih ada satu klik manusia

Satu-satunya pekerjaan manusia yang tersisa adalah satu klik setuju di Telegram sebelum publish. Itu keputusan, bukan sisa yang belum sempat diotomasi.

Alasannya asimetri biaya. Satu klik memakan lima detik dan bisa dilakukan dari HP. Satu kesalahan yang terlanjur publik tidak berhenti salah setelah videonya dihapus: sudah masuk rekomendasi, sudah ditonton, sudah di-screenshot. Untuk konten finansial, kesalahan yang paling mungkin lolos justru yang paling mahal — angka yang salah di layar. Prompt-nya melarang keras angka yang tidak berasal dari narasi, tapi larangan di prompt itu kecenderungan, bukan jaminan, dan tidak ada satu pun tahap otomatis di pipeline ini yang bisa membedakan angka benar dari angka yang terdengar benar.

Gerbangnya juga sengaja tidak galak. Ia menerima dua status: done dan needs_review. Render yang skor kualitasnya jelek tetap dikirim ke Telegram untuk dilihat manusia, bukan dibuang otomatis, karena penilaian mesin soal 'jelek' sering salah ke dua arah. Yang benar-benar dihentikan hanya kegagalan keras: status di luar dua itu, atau durasi di bawah 20 detik.

Kalau kamu membangun ulang sistem seperti ini, tahap approve adalah bagian pertama yang dibuat, bukan terakhir. Pipeline tanpa gerbang bukan pipeline yang lebih cepat — itu pipeline yang kesalahannya keluar duluan.

Yang dicoba lalu dibuang

Bagian ini yang membuat bagian lain layak dipercaya. Semua yang di bawah ini sempat jalan di produksi sebelum dibuang.

Suara Indonesia lewat mesin TTS yang tidak punya suara Indonesia. Waktu itu enginenya dipilih karena murah dan sudah terpasang. Pengecekan langsung ke daftar suaranya: 260 suara, 16 kode bahasa, nol di antaranya Indonesia. Yang sebenarnya terjadi adalah suara Amerika membaca teks Indonesia secara cross-lingual, dan itu terdengar persis seperti kedengarannya. Dibuang — sebelum itu, tidak ada satu pun baris kode yang salah.

Suara super murah yang hasilnya kaku. Pengganti berikutnya sekitar 28 kali lebih murah dari ElevenLabs multilingual, dan secara angka itu keputusan yang jelas. Setelah dipakai memproduksi antrean, hasilnya kaku dan tidak enak didengar sepanjang 10 menit. Antrean 84 video hasil suara itu tidak diteruskan, dan pipeline balik ke suara kloning ElevenLabs. Penghematan yang menghasilkan video yang tidak ditonton bukan penghematan.

Jadwal harian otomatis. WF6 dulu punya trigger terjadwal; sekarang on-demand saja. Dua alasan: klip film masuk kuota langganan ~900 klip/bulan yang fail-closed kalau habis, dan produksi terjadwal berarti kuota itu terbakar oleh topik yang dipilih kalender, bukan oleh topik yang memang layak. Volume gampang; volume yang layak ditonton tidak.

Rencana perbaikan yang cuma menyentuh config. Ada satu putaran tuning yang isinya mengubah angka-angka konfigurasi supaya video lebih panjang dan lebih variatif. Hasilnya hampir nol, karena jumlah beat film sebenarnya ditentukan oleh berapa klip yang berhasil diunduh, bukan oleh angka di config. Perbaikan yang benar ada di pengunduhnya. Pelajarannya: sebelum menulis rencana, cari dulu di mana angka itu benar-benar diputuskan.

Mesin terkecil yang masuk akal dibangun satu weekend

Jangan mulai dari sini. Sistem di halaman ini tumbuh berbulan-bulan dan sebagian besar isinya adalah tambalan untuk masalah yang belum akan kamu punya. Yang perlu kamu buktikan dulu cuma satu: ada satu potong pekerjaan berulang yang selesai tanpa kamu, dan kamu tetap pegang kendali di ujungnya.

Bentuk terkecilnya empat bagian. Satu trigger: webhook n8n yang menerima satu field teks. Satu panggilan AI: HTTP request ke API model, dengan system prompt yang kamu tulis sendiri, max_tokens diisi eksplisit, dan mode thinking dimatikan kalau kamu minta output terstruktur. Satu output yang disimpan ke tempat yang bisa kamu buka lagi — baris database atau file, bukan cuma pesan chat. Satu titik approve: pesan Telegram dengan dua tombol URL yang memanggil webhook kedua, membawa id pekerjaan dan aksi (approve atau reject) di query string; webhook itu cuma mengubah status baris tadi.

Output pertamamu jangan video. Video menyeret masuk render, encoding, subtitle, musik, dan hak pakai — lima sistem sekaligus, dan tidak satu pun mengajari kamu hal yang belum kamu tahu tentang otomasi. Naskah, ringkasan, draft caption, balasan email: semuanya membuktikan bentuk yang sama dengan sepersepuluh permukaan gagal.

Sebelum dianggap selesai, tambahkan satu hal yang biasanya ditinggal: alarm untuk hasil nol. Mode gagal yang paling mahal di pipeline seperti ini bukan crash — crash itu berisik dan langsung kamu tahu. Yang mahal adalah eksekusi hijau yang mengembalikan nol item, tiap hari, selama berminggu-minggu. Kirim pesan ketika hasilnya nol, sama seperti kamu mengirim pesan ketika hasilnya jadi.

Kalau itu bertahan seminggu tanpa kamu sentuh, baru tambah tahap kedua. Setiap tahap baru menambah satu cara baru untuk gagal diam-diam, dan enam jebakan di atas semuanya jenis itu.

Perubahan

  • 22 Juli 2026Versi pertama.

Ambil kit-nya

Isi halaman ini gratis dan lengkap — kit ini cuma buat yang mau langsung eksekusi: workflow JSON yang sudah disanitasi, file prompt, .env.example, checklist deploy, dan kalkulator biaya.

Cuma dipakai buat ngirim kit ini dan kabar kalau blueprint-nya diperbarui. Nggak dijual, nggak dibagi ke siapa pun.

Mau lihat sistemnya jalan, bukan cuma bacanya? Playlist “Uang di Layar” isinya video yang diproduksi pakai pipeline ini.

Nggak mau ngurusin sendiri?

Semua yang ada di halaman ini bisa dibangun sendiri — memang itu tujuannya ditulis. Tapi kalau yang kamu butuhkan bukan belajar n8n, melainkan sistemnya jalan minggu depan untuk bisnismu, itu yang saya kerjakan.

Lihat jasa otomatisasi

Have a project in mind?