Meaning before speed
Names, numbers, amounts, percentages, negations, claims and instructions should survive timing adaptation. AI is used to rewrite long dialogue more efficiently, not to delete facts.
Gemini YouTube Dubber turns public YouTube content or local video into translated speech, subtitles and a synchronized dubbing package. Its core idea is simple: preserve meaning, preserve the source timeline, and never “solve” timing by making speech unnaturally fast or slow.
Public YouTube URL / local video
↓
Gemini video understanding + translation
↓
Semantic source units + speaker/timestamp locks
↓
AI Timing Director (pre-TTS compression only when needed)
↓
Edge / Gemini speech synthesis
↓
Measured Timing Feedback (hard natural-rate limit: 1.10×)
↓
FFmpeg timeline placement + silence + soundtrack mix
↓
WAV + SRT + transcript + manifest → final MP4Normal translation is not enough for dubbing. A correct translated sentence may take 30–70% longer to speak than the original. If software blindly stretches or speeds audio to fit, the result sounds robotic. This project treats translation, timing, speech generation and timeline placement as one coordinated system.
Names, numbers, amounts, percentages, negations, claims and instructions should survive timing adaptation. AI is used to rewrite long dialogue more efficiently, not to delete facts.
Short speech is accepted as short. Remaining room stays silent. Overlong speech is rewritten and regenerated instead of being aggressively time-stretched.
Speech stays anchored to the original onset and neighboring cues. Conservative semantic merges remove artificial micro-deadlines without moving the whole timeline.
Gemini analyzes the public video URL or local media, producing speaker-aware timestamps, transcription and target-language translation. A reusable checkpoint prevents repeated video analysis when the same source is processed again.
Adjacent same-speaker fragments can be conservatively merged when they are clearly one unfinished sentence. v0.5.10 also bridges an ultra-short heading across a small real pause when the next same-speaker cue is the obvious longer continuation.
Before TTS, translated text is evaluated against its source slot. Only text predicted to be too long is compressed. Cloud mode currently targets about 72% pre-TTS occupancy and processes timing work in batches of ten.
Cloud validation uses Edge Neural TTS as the primary speech engine, while the project also contains Gemini TTS support and fallback orchestration. The project does not clone the source voice.
Generated audio duration is measured. If it would require more than the hard 1.10× speed ceiling, Gemini receives the exact ratio and rewrites the text again. Up to three measured feedback passes are allowed.
FFmpeg places speech at source positions, preserves silence, performs required audio operations and creates the dubbing track. Local finalization can combine the package with the source video into MP4.
| Layer | Current role | Engineering choice |
|---|---|---|
| AI understanding / timing | Transcription, translation, semantic compression and timing feedback | gemini-3.5-flash-lite in the current Cloud workflow; model names remain environment-configurable. |
| Speech | Dubbing synthesis | Edge Neural TTS is primary in the validated Cloud path; Gemini TTS support remains available in the codebase. |
| Synchronization | Source-timestamp alignment | segment_locked / onset-locked design, semantic continuation merging, micro-cue bridge, real-silence borrowing ≤ 1.50 s. |
| Timing control | Avoid rushed or slow-motion voice | Pre-TTS Timing Director + measured post-TTS feedback; hard speed ceiling 1.10×; short speech is never expanded just to fill time. |
| Media | Audio timing, mixing and final mux | FFmpeg via imageio-ffmpeg; no pydub/audioop dependency, keeping Python 3.13 compatibility. |
| Local UI | Interactive desktop/browser workflow | Streamlit, plus CLI and Windows/macOS/Linux launch paths. |
| Cloud | Remote processing | GitHub Actions with encrypted GEMINI_API_KEY, reusable transcript/TTS cache and downloadable artifacts. |
| Testing | Regression protection | CI matrix for Python 3.11, 3.12 and 3.13, including timing, API-shape and cloud-workflow tests. |
A subtitle boundary is not always a speech boundary. The system can merge only carefully defined same-speaker continuations, removing impossible 1–2 second deadlines while preserving the first onset and final end time.
GitHub runner IPs can be blocked by YouTube media delivery. The cloud workflow therefore lets Gemini understand the public URL and builds audio/SRT/JSON remotely, while final source-video download and mux can be completed locally.
The important milestone is not a demo screenshot; it is a completed end-to-end Cloud Dub package after repeated real failures exposed timing, quota, model-availability and micro-cue edge cases.
The v0.5.10 semantic micro-cue fix reduced the source to 37 timestamp-locked speech units and completed all 37. Difficult chunks were measured, AI-compressed and re-synthesized while the 1.10× hard limit remained unchanged. The workflow then built and uploaded the complete cloud dubbing package.
The project survived API-format changes, Python 3.13 audio dependency issues, 429 quota exhaustion, 503 capacity failures, retired model names and timing-convergence failures. Each failure became a regression rule or architectural change.
Successful packages include dubbed_audio.wav, dubbed.srt, transcript.json, manifest.json and timing diagnostic files so the process is inspectable rather than opaque.
Persian is the default target, but the pipeline accepts other target languages and can be driven locally, by CLI, Docker or GitHub Actions.
Synchronization is timestamp/utterance based. It can sound naturally aligned, but it does not yet move a face or solve viseme-level mouth matching.
The cloud package is complete for dubbing audio and metadata, but YouTube bot protection can prevent GitHub from reliably downloading the source video. Final MP4 mux may therefore happen on the user machine.
Gemini availability, free-tier limits and model names are external constraints. The code is configurable and quota-aware, but it cannot create provider capacity.
The current project uses prebuilt voices. This is safer and simpler, but it does not reproduce each original speaker's identity or emotion exactly.
These are engineering opportunities, not promises. The highest-value work is the work that makes quality measurable, reduces repeated API cost and removes the final manual step.
Cache successful AI timing rewrites by source text, target language, voice and timing slot so the same hard chunk never consumes Gemini quota twice.
Add reliable user-upload/object-storage ingestion or another authorized media source so GitHub can mux the final MP4 without depending on blocked YouTube runner downloads.
Back-transcribe generated audio and score semantic fidelity, names/numbers, timing, loudness, clipping and missing speech before an artifact is declared successful.
Provide a visual timeline where users can edit translation, merge/split cues, lock terminology, choose a voice per speaker and regenerate only selected regions.
Add project dictionaries for names, brands, technical vocabulary, units and preferred translations, with explicit protection for numbers and negation.
Go beyond two voice roles with better diarization, speaker profiles and consistent voice assignment across long videos and episode series.
Introduce voice similarity or cloning only with explicit consent, provenance records and clear disclosure, while retaining prebuilt voices as the safe default.
Generate phoneme or viseme timing and optionally connect to a lip-sync model for close-up talking-head content where mouth alignment matters.
Add a dashboard for multiple jobs, cost/request counters, cache hit rate, provider latency, failure classification and per-language quality history.
Package Windows, macOS and Linux binaries so non-Python users can run the local workflow, update safely and finalize cloud packages with one click.
Gemini YouTube Dubber, herkese açık YouTube içeriklerini veya yerel videoları çevrilmiş konuşma, altyazı ve senkronize dublaj paketine dönüştürür. Temel prensip nettir: anlamı koru, kaynak zaman çizgisini koru ve zamanlamayı sesi yapay biçimde aşırı hızlandırarak ya da yavaşlatarak “çözme”.
YouTube URL / yerel video
↓
Gemini video anlama + çeviri
↓
Anlamsal konuşma birimleri + konuşmacı/zaman kilitleri
↓
AI Timing Director
↓
Edge / Gemini konuşma sentezi
↓
Ölçülmüş Timing Feedback (sert sınır: 1.10×)
↓
FFmpeg zaman çizgisi yerleşimi + sessizlik + miks
↓
WAV + SRT + transcript + manifest → son MP4Dublaj için yalnızca doğru çeviri yeterli değildir. Doğru çevrilmiş bir cümlenin okunması orijinalden çok daha uzun sürebilir. Ses körlemesine sıkıştırılırsa robotik bir sonuç çıkar. Bu proje çeviri, süre, konuşma üretimi ve zaman çizgisi yerleşimini tek bir sistem olarak tasarlar.
İsimler, sayılar, tutarlar, yüzdeler, olumsuzluklar, iddialar ve talimatlar korunmalıdır. AI, gerçekleri silmek için değil, uzun ifadeyi daha verimli söylemek için kullanılır.
Kısa konuşma kısa kalır; kalan süre sessizliktir. Uzun konuşma ise aşırı time-stretch yerine yeniden yazılır ve yeniden üretilir.
Konuşma başlangıçları kaynak videoya bağlı kalır. Yalnızca güvenli anlamsal birleştirmeler yapay mikro-deadline'ları kaldırır.
Gemini konuşmacı etiketleri, zaman damgaları, transkript ve hedef dil çevirisi üretir. Checkpoint sistemi aynı videonun tekrar analiz edilmesini önler.
Aynı konuşmacının bölünmüş ama tek cümle olan parçaları kontrollü biçimde birleştirilebilir. v0.5.10, çok kısa bir başlığı küçük bir gerçek sessizlik üzerinden açıkça devam eden sonraki cümleyle birleştirebilir.
TTS'den önce metin kaynak konuşma yuvasıyla karşılaştırılır. Yalnızca fazla uzun görünen metin kısaltılır. Bulut ayarı yaklaşık %72 doluluk hedefi ve 10'luk batch kullanır.
Doğrulanmış bulut yolunda Edge Neural TTS birincil motordur. Kod tabanında Gemini TTS ve fallback orkestrasyonu da vardır. Kaynak ses klonlanmaz.
Gerçek ses süresi ölçülür. 1.10×'ten fazla hız gerekecekse Gemini'ye tam oran geri verilir ve metin tekrar sıkıştırılır. En fazla üç geçiş vardır.
FFmpeg konuşmaları kaynak konumlarına yerleştirir, sessizliği korur, ses işlemlerini uygular ve dublaj kanalını oluşturur. Yerel finalize adımı kaynak videoyla MP4 mux yapabilir.
| Katman | Görev | Teknik seçim |
|---|---|---|
| AI anlama / timing | Transkripsiyon, çeviri, anlamsal sıkıştırma, timing feedback | Bulut workflow'unda gemini-3.5-flash-lite; model adları environment variable ile değiştirilebilir. |
| Konuşma | Dublaj sentezi | Doğrulanmış bulut yolunda Edge Neural TTS; kodda Gemini TTS desteği de bulunur. |
| Senkronizasyon | Kaynak zaman damgalarına hizalama | segment_locked, onset-lock, Semantic Lock, Micro-Cue Bridge, ≤1.50 sn sessizlik ödünç alma. |
| Timing kontrolü | Aşırı hızlı/yavaş sesi önleme | Pre-TTS Timing Director + ölçümlü post-TTS feedback; sert sınır 1.10×; kısa ses slot doldurmak için uzatılmaz. |
| Medya | Timing, miks, mux | FFmpeg + imageio-ffmpeg; pydub/audioop yok; Python 3.13 uyumlu. |
| Arayüz | Yerel kullanım | Streamlit + CLI + Windows/macOS/Linux çalıştırma yolları. |
| Bulut | Uzaktan dublaj | GitHub Actions, şifreli GEMINI_API_KEY, transcript/TTS cache ve artifact. |
| Test | Regresyon koruması | Python 3.11/3.12/3.13 CI matrisi ve timing/API/cloud testleri. |
Altyazı sınırı her zaman gerçek konuşma sınırı değildir. Sistem yalnızca net kurallara uyan aynı konuşmacı devamlarını birleştirir ve ilk başlangıç ile son bitişi korur.
GitHub runner IP'leri YouTube tarafından engellenebilir. Bu yüzden bulut URL'yi Gemini'ye analiz ettirir ve ses/SRT/JSON üretir; son video indirme ve mux yerelde yapılabilir.
Önemli kilometre taşı bir ekran görüntüsü değil; gerçek 429, 503, model kaldırılması ve timing hatalarından sonra başarıyla tamamlanan bir Cloud Dub paketidir.
v0.5.10 Micro-Cue düzeltmesi kaynağı 37 timestamp-kilitli konuşma birimine dönüştürdü ve 37'sinin tamamı işlendi. Zor parçalar ölçüldü, AI ile kısaltıldı ve tekrar sentezlendi; 1.10× üst sınırı değiştirilmedi.
API format değişiklikleri, Python 3.13 bağımlılık sorunu, quota, 503 kapasite, kullanımdan kalkan modeller ve timing hataları mimari kurala veya regresyon testine dönüştürüldü.
dubbed_audio.wav, dubbed.srt, transcript.json, manifest.json ve timing raporları süreci şeffaf tutar.
Varsayılan hedef Farsçadır; ancak farklı hedef diller, yerel UI, CLI, Docker ve GitHub Actions yolları desteklenir.
Senkronizasyon konuşma/timestamp düzeyindedir; ağız şekli/viseme üretmez.
Bulut ses ve metadata paketini tamamlar; fakat GitHub'ın YouTube indirmesi bot korumasına takılabildiği için son MP4 mux yerelde gerekebilir.
Model adları, free-tier sınırları ve kapasite Google tarafından değiştirilebilir. Kod uyarlanabilir ama sağlayıcı kapasitesi yaratamaz.
Hazır sesler kullanılır; her orijinal konuşmacının kimliğini birebir taklit etmez.
Bunlar teknik geliştirme önerileridir. En yüksek değer: tekrar API tüketimini azaltmak, kaliteyi ölçmek ve son manuel adımı ortadan kaldırmaktır.
Başarılı sıkıştırmaları kaynak metin + hedef dil + ses + slot anahtarıyla saklayarak aynı zor cümlenin tekrar Gemini isteği tüketmesini engelle.
Kullanıcı yükleme/object storage gibi güvenilir bir kaynak ekleyerek GitHub'ın YouTube indirmesine bağımlı olmadan final mux yap.
Üretilen sesi tekrar ASR ile çözümleyip anlam, isim/sayı, timing, loudness, clipping ve eksik konuşmayı puanla.
Çeviri düzenleme, cue birleştirme/bölme, terminoloji kilidi, konuşmacı sesi ve seçili bölge regenerasyonu sun.
İsim, marka, teknik terim, birim ve tercih edilen çeviri sözlükleri ekle; sayı ve olumsuzlukları özel olarak kilitle.
İki rolün ötesinde diarization, konuşmacı profili ve uzun videolarda tutarlı ses eşleştirme geliştir.
Açık rıza, provenance ve açıklama şartıyla opsiyonel ses benzerliği/klonlama; hazır sesler güvenli varsayılan olarak kalsın.
Yakın plan konuşma videoları için fonem/viseme timing ve isteğe bağlı lip-sync modeli entegrasyonu.
Python kurmadan çalışan Windows, macOS ve Linux binary'leri ve tek tıkla cloud finalize akışı.
Çoklu iş kuyruğu, quota/cost sayaçları, cache hit oranı, latency, hata sınıflandırması ve kalite geçmişi.
Gemini YouTube Dubber transforma vídeos públicos de YouTube o archivos locales en voz traducida, subtítulos y un paquete de doblaje sincronizado. La idea central es conservar el significado, conservar la línea de tiempo original y no “resolver” la duración acelerando o ralentizando la voz de forma artificial.
URL pública / vídeo local
↓
Comprensión de vídeo + traducción con Gemini
↓
Unidades semánticas + bloqueos de hablante/tiempo
↓
AI Timing Director
↓
Síntesis de voz Edge / Gemini
↓
Timing Feedback medido (límite duro: 1.10×)
↓
FFmpeg: colocación + silencios + mezcla
↓
WAV + SRT + transcript + manifest → MP4 finalTraducir correctamente no basta para doblar. Una frase traducida puede necesitar mucho más tiempo que la original. Si el software fuerza el audio a entrar mediante time-stretch agresivo, la voz suena robótica. Este proyecto coordina traducción, duración, síntesis y montaje temporal como un único sistema.
Nombres, números, importes, porcentajes, negaciones, afirmaciones e instrucciones deben conservarse. La IA compacta la formulación, no los hechos.
Si la voz es corta, se acepta corta y el resto queda en silencio. Si es demasiado larga, el texto se reescribe y se vuelve a sintetizar.
Los inicios permanecen anclados al vídeo original. Fusiones semánticas conservadoras eliminan micro-plazos imposibles sin desplazar toda la línea de tiempo.
Gemini genera transcripción, hablantes, timestamps y traducción. Un checkpoint reutilizable evita analizar de nuevo el mismo vídeo.
Fragmentos contiguos del mismo hablante pueden fusionarse cuando forman claramente una sola frase. v0.5.10 permite además unir un encabezado ultracorto con su continuación tras una pequeña pausa real.
Antes del TTS, el texto se compara con su slot de origen. Solo se comprime si parece demasiado largo. En cloud se usa un objetivo previo cercano al 72% y lotes de 10.
La ruta cloud validada usa Edge Neural TTS como motor principal. El código mantiene soporte para Gemini TTS. No clona la identidad vocal original.
Se mide la duración real. Si exigiría más de 1.10×, Gemini recibe el ratio exacto y vuelve a compactar el texto. Máximo: tres pasadas.
FFmpeg coloca cada voz, conserva silencios y construye la pista final. La finalización local puede muxearla con el vídeo fuente.
| Capa | Función | Elección técnica |
|---|---|---|
| IA / timing | Transcripción, traducción, compresión semántica y feedback | gemini-3.5-flash-lite en el workflow cloud actual; modelos configurables por entorno. |
| Voz | Síntesis de doblaje | Edge Neural TTS en la ruta cloud validada; soporte Gemini TTS en el código. |
| Sincronización | Alineación con timestamps | segment_locked, onset-lock, Semantic Lock, Micro-Cue Bridge y préstamo de silencio ≤ 1.50 s. |
| Control temporal | Evitar voz acelerada/lenta | Timing Director pre-TTS + feedback post-TTS medido; techo 1.10×; la voz corta no se expande para rellenar. |
| Media | Timing, mezcla y mux | FFmpeg + imageio-ffmpeg, sin pydub/audioop, compatible con Python 3.13. |
| UI local | Uso interactivo | Streamlit + CLI + launchers multiplataforma. |
| Cloud | Procesamiento remoto | GitHub Actions, secreto GEMINI_API_KEY, caché reutilizable y artifacts. |
| Pruebas | Protección contra regresiones | Matriz CI Python 3.11/3.12/3.13 y tests de timing, API y workflow. |
Un corte de subtítulos no siempre es un corte real del habla. El sistema fusiona solo continuaciones muy bien definidas y conserva el primer inicio y el último final.
YouTube puede bloquear IPs de runners de GitHub. Por eso Gemini puede comprender la URL pública y el cloud genera audio/SRT/JSON; la descarga final del vídeo y el mux pueden hacerse localmente.
El hito importante no es una captura de pantalla: es un paquete Cloud Dub completo después de convertir fallos reales de cuota, capacidad, modelos retirados y timing en cambios de arquitectura.
La corrección Micro-Cue de v0.5.10 produjo 37 unidades bloqueadas por timestamp y completó las 37. Los chunks difíciles se midieron, compactaron con IA y resintetizaron sin elevar el techo duro de 1.10×.
Cambios de formato API, Python 3.13, 429, 503, modelos retirados y fallos de convergencia terminaron convertidos en reglas o tests.
El paquete contiene dubbed_audio.wav, dubbed.srt, transcript.json, manifest.json y diagnósticos de timing.
Persa es el destino por defecto, pero el pipeline admite otros idiomas y puede ejecutarse mediante UI local, CLI, Docker o GitHub Actions.
La alineación es por utterance/timestamp, no por visemas o movimiento de boca.
El audio y metadata se producen en cloud; la protección anti-bot de YouTube puede obligar a hacer el mux MP4 localmente.
Google puede cambiar límites, capacidad y nombres de modelos. El software se adapta, pero no controla la capacidad del proveedor.
Usa voces preconstruidas; no pretende copiar exactamente la identidad de cada hablante.
Son oportunidades de ingeniería, no promesas. El mayor valor está en medir la calidad, reducir llamadas repetidas y eliminar el último paso manual.
Guardar compresiones exitosas por texto, idioma, voz y slot para no gastar cuota dos veces en el mismo chunk difícil.
Añadir ingestión fiable por upload/object storage para hacer el mux cloud sin depender de descargas de YouTube desde runners.
Retranscribir el audio generado y puntuar fidelidad semántica, nombres/números, timing, loudness, clipping y voz ausente.
Editar traducción, fusionar/dividir cues, bloquear terminología, asignar voces y regenerar solo regiones seleccionadas.
Diccionarios de nombres, marcas, términos técnicos, unidades y traducciones preferidas, con protección explícita de números y negaciones.
Mejor diarización y asignación consistente de más de dos perfiles de voz en vídeos largos o series.
Similitud/clonación opcional solo con consentimiento, provenance y divulgación clara; voces preconstruidas como opción segura.
Timing fonético/visema y conexión opcional con un modelo de lip-sync para primeros planos.
Binarios Windows/macOS/Linux sin instalación de Python y finalización cloud con un clic.
Jobs múltiples, contadores de cuota/coste, cache hit rate, latencia, clasificación de fallos e historial de calidad.