How It Works — Step by Step
Step 1: Upload
Upload an audio or video file (MP3, WAV, MP4, AVI, MKV, MOV, WEBM). Maximum 60 seconds and 400 MB. The app transcribes the English speech and identifies each speaker automatically.
Step 2: Edit Segments
Review the transcription. You can:
- Edit text — fix any transcription mistakes directly in the table.
- Auto Translate — translate all unlocked lines to Arabic using AI.
- Add Tashkeel — add Arabic diacritics (harakat) to the translated text.
- Detect Emotions — AI listens to each segment's voice tone and assigns a speaking style.
- Auto-Fix Timing — re-sync segment timestamps to match the original audio precisely.
- 🔒 Lock — protect specific lines from being changed by Auto-Fix, Translate, or Tashkeel.
Step 3: Voice Cloning Setup
The app analyzes how much speech each speaker has. Speakers with enough audio (≥1 second) can be cloned — the AI copies their unique voice from the video so the Arabic dub sounds like them.
Step 3.5: Choose Speakers to Clone
Review the quality guidance for each speaker. Uncheck any speaker you don't want to clone. Speakers with very little audio will produce poor clones — use a studio library voice instead.
Step 4: Assign Voices
Pick a voice for each speaker. You can use:
- Cloned voices — copied from your video (if cloning succeeded).
- Studio library voices — high-quality professional voices (numbered, names hidden).
Use 🎲 Auto-Assign to let the app pick the best match automatically. Two speakers never share the same numbered voice.
Step 5: Generate Arabic Audio
Click Generate Arabic Audio. The AI speaks each line in Arabic using the assigned voice and emotion/style. This may take 1–3 minutes depending on the clip length.
Step 6: Review & Fine-Tune
Listen to the result. You can:
- 🔄 Re-speak a single line — regenerate just one line without redoing the whole clip.
- 🎚️ Fine-Tune Timeline — drag segment blocks left/right (in milliseconds) to perfectly align with lip movements. Blocks cannot overlap.
- 🎬 Merge into Video — combine the dubbed audio with the original video and background music.
- 💾 Save / Load Project — save your work at any time as a JSON file. Load it later to continue editing without re-transcribing.
Frequently Asked Questions
What file formats are supported?
Audio: MP3, WAV, FLAC, AAC, OGG. Video: MP4, AVI, MKV, MOV, WEBM. Maximum duration: 60 seconds. Maximum file size: 400 MB.
How does voice cloning work?
The AI extracts each speaker's voice characteristics from the original audio and creates a synthetic copy. This copy is then used to speak the Arabic translation in a voice that sounds similar to the original speaker. At least 1 second of clear speech is required; 10+ seconds produces the best results.
What are credits and how are costs calculated?
Credits are our internal billing unit: 1 credit = $0.01 USD (100 credits = $1.00). Translation and analysis cost a few cents per clip. Voice generation is charged by character count. Every button shows its estimated cost before you click it, and actual usage appears in the green box in Step 1.
What does the 🔒 lock do?
Locking a line protects it from three actions: Auto-Fix Timing, Auto Translate, and Add Tashkeel. Locked lines keep their current text, timing, and formatting untouched. Use this for lines you've manually perfected.
Can I combine multiple emotion/style tags?
Yes! Use the dropdown to add tags one at a time, or type directly in the text box (e.g., "confident, calm"). Only recognized tags from the official list are accepted; unrecognized words are removed automatically. Examples: "anxious, afraid", "playful, teasing", "calm, firm".
Why does my cloned voice sound robotic?
This usually means the speaker didn't have enough clear audio in the original clip. Check the quality guidance in Step 3.5 — if it says ⚠️ or ❌, switch that speaker to a studio library voice in Step 4 for better results.
How does the timeline fine-tuning work?
In Step 6, click 🎚️ Fine-Tune Timeline. Each blue block represents one spoken line. Drag blocks left or right to shift their timing by milliseconds. A time ruler at the top shows your position. Blocks stop at adjacent segments to prevent overlap. Click ✅ Confirm to rebuild the audio with your adjustments.
Does the merged video keep the original background music?
Yes. The app separates vocals from background audio during processing. When you merge, the dubbed Arabic vocals are mixed with the original background music/sound effects at balanced volumes.
Can I save my progress and come back later?
Yes. Click 💾 Save Project in Step 2 to download a JSON file with all your segments, translations, voice assignments, and settings. Later, click 📂 Load Project to restore everything exactly where you left off.
Is lip-sync available?
Not yet. Lip-sync (matching lip movements to the new language) is planned for a future update. Currently, the app produces high-quality dubbed audio that you can merge with your video.
Will the resolution or quality of my video be affected?
No. The original video stream is copied unchanged; only the audio track is replaced (or mixed) with the dubbed Arabic audio. Resolution, bitrate, frame rate, and visual quality stay exactly as your source file.
Are my uploaded files and generated videos saved on your servers?
No. Your uploaded files, transcriptions, and generated audio/video are stored temporarily on the server only while your session is active. They are automatically deleted when the server restarts or after a short period. You must download your results during your session — once you close the browser or the server recycles, the files are gone permanently. We do not keep copies of your media. For your safety, always click the download buttons in Step 6 before leaving the page.
كيف يعمل التطبيق — خطوة بخطوة
الخطوة ١: رفع الملف
ارفع ملف صوتي أو مرئي (MP3، WAV، MP4، AVI، MKV، MOV، WEBM). الحد الأقصى ٦٠ ثانية و ٤٠٠ ميغابايت. يقوم التطبيق بنسخ الكلام الإنجليزي وتحديد كل متحدث تلقائياً.
الخطوة ٢: تحرير المقاطع
راجع النسخ المكتوبة. يمكنك:
- تعديل النص — أصلح أي أخطاء في النسخ مباشرة من الجدول.
- الترجمة التلقائية — ترجم جميع الأسطر غير المقفلة إلى العربية باستخدام الذكاء الاصطناعي.
- إضافة التشكيل — أضف الحركات العربية إلى النص المترجم.
- كشف المشاعر — يستمع الذكاء الاصطناعي لنبرة صوت كل مقطع ويحدد أسلوب التحدث المناسب.
- إصلاح التوقيت التلقائي — أعد مزامنة توقيت المقاطع مع الصوت الأصلي بدقة.
- 🔒 قفل — احمِ أسطراً محددة من التغيير بواسطة إصلاح التوقيت أو الترجمة أو التشكيل.
الخطوة ٣: إعداد استنساخ الأصوات
يحلل التطبيق كمية الكلام المتوفرة لكل متحدث. المتحدثون الذين لديهم صوت كافٍ (≥ ثانية واحدة) يمكن استنساخ أصواتهم — ينسخ الذكاء الاصطناعي خصائص صوتهم الفريدة من الفيديو ليبدو الدبلج العربي مشابهاً لهم.
الخطوة ٣.٥: اختيار المتحدثين للاستنساخ
راجع إرشادات الجودة لكل متحدث. ألغِ تحديد أي متحدث لا تريد استنساخه. المتحدثون ذوو الصوت القليل جداً سيُنتجون نسخاً ضعيفة — استخدم صوتاً من مكتبة الاستوديو بدلاً من ذلك.
الخطوة ٤: تعيين الأصوات
اختر صوتاً لكل متحدث. يمكنك استخدام:
- الأصوات المستنسخة — منسوخة من الفيديو الخاص بك (إذا نجح الاستنساخ).
- أصوات مكتبة الاستوديو — أصوات احترافية عالية الجودة (مُرقّمة، الأسماء مخفية).
استخدم 🎲 تعيين تلقائي ليدع التطبيق يختار أفضل تطابق تلقائياً. لا يتشارك متحدثان نفس الصوت المُرقّم أبداً.
الخطوة ٥: توليد الصوت العربي
اضغط توليد الصوت العربي. ينطق الذكاء الاصطناعي كل سطر بالعربية باستخدام الصوت والمشاعر/الأسلوب المعين. قد يستغرق هذا ١–٣ دقائق حسب طول المقطع.
الخطوة ٦: المراجعة والضبط الدقيق
استمع للنتيجة. يمكنك:
- 🔄 إعادة نطق سطر واحد — أعد توليد سطر واحد فقط دون إعادة المقطع كاملاً.
- 🎚️ ضبط الجدول الزمني — اسحب مربعات المقاطع يميناً/يساراً (بالمللي ثانية) لمحاذاة حركة الشفاه تماماً. لا يمكن للمربعات أن تتداخل.
- 🎬 دمج مع الفيديو — ادمج الصوت المدبلج مع الفيديو الأصلي والموسيقى الخلفية.
- 💾 حفظ / تحميل المشروع — احفظ عملك في أي وقت كملف JSON. حمّله لاحقاً لمتابعة التحرير دون إعادة نسخ.
الأسئلة الشائعة
ما هي صيغ الملفات المدعومة؟
الصوت: MP3، WAV، FLAC، AAC، OGG. الفيديو: MP4، AVI، MKV، MOV، WEBM. الحد الأقصى للمدة: ٦٠ ثانية. الحد الأقصى لحجم الملف: ٤٠٠ ميغابايت.
كيف يعمل استنساخ الأصوات؟
يستخرج الذكاء الاصطناعي خصائص صوت كل متحدث من الصوت الأصلي وينشئ نسخة اصطناعية. تُستخدم هذه النسخة لنطق الترجمة العربية بصوت مشابه للمتحدث الأصلي. مطلوب ثانية واحدة على الأقل من كلام واضح؛ ١٠ ثوانٍ فأكثر تعطي أفضل النتائج.
ما هي النقاط وكيف تُحسب التكاليف؟
النقاط هي وحدة الفوترة الداخلية: نقطة واحدة = ٠.٠١ دولار أمريكي (١٠٠ نقطة = دولار واحد). الترجمة والتحليل تكلف بضعة سنتات لكل مقطع. توليد الصوت يُحسب بعدد الأحرف. كل زر يعرض تكلفته التقديرية قبل الضغط عليه، والاستخدام الفعلي يظهر في المربع الأخضر في الخطوة ١.
ماذا يفعل القفل 🔒؟
قفل سطر يحميه من ثلاث عمليات: إصلاح التوقيت التلقائي، الترجمة التلقائية، وإضافة التشكيل. الأسطر المقفلة تحتفظ بالنص والتوقيت والتنسيق الحالي دون تغيير. استخدم هذا للأسطر التي أتقنتها يدوياً.
هل يمكنني دمج عدة وسوم للمشاعر/الأسلوب؟
نعم! استخدم القائمة المنسدلة لإضافة الوسوم واحداً تلو الآخر، أو اكتب مباشرة في مربع النص (مثال: "واثق، هادئ"). فقط الوسوم المعتمدة من القائمة الرسمية مقبولة؛ الكلمات غير المعروفة تُحذف تلقائياً. أمثلة: "قلق، خائف"، "مرح، مازح"، "هادئ، حازم".
لماذا يبدو صوتي المستنسخ آلياً؟
هذا يعني عادةً أن المتحدث لم يكن لديه صوت كافٍ في المقطع الأصلي. تحقق من إرشادات الجودة في الخطوة ٣.٥ — إذا كانت تقول ⚠️ أو ، غيّر صوت هذا المتحدث إلى صوت مكتبة الاستوديو في الخطوة ٤ للحصول على نتائج أفضل.
كيف يعمل الضبط الدقيق للجدول الزمني؟
في الخطوة ٦، اضغط 🎚️ ضبط الجدول الزمني. كل مربع أزرق يمثل سطراً منطوقاً واحداً. اسحب المربعات يميناً أو يساراً لتغيير توقيتها بالمللي ثانية. مسطرة زمنية في الأعلى تظهر موقعك. تتوقف المربعات عند المقاطع المجاورة لمنع التداخل. اضغط ✅ تأكيد لإعادة بناء الصوت بتعديلاتك.
هل يحتفظ الفيديو المدمج بالموسيقى الخلفية الأصلية؟
نعم. يفصل التطبيق الأصوات عن الخلفية الصوتية أثناء المعالجة. عند الدمج، تُخلط الأصوات العربية المدبلجة مع الموسيقى/المؤثرات الخلفية الأصلية بمستويات صوت متوازنة.
هل يمكنني حفظ تقدمي والعودة لاحقاً؟
نعم. اضغط 💾 حفظ المشروع في الخطوة ٢ لتنزيل ملف JSON يحتوي على جميع المقاطع والترجمات وتعيينات الأصوات والإعدادات. لاحقاً، اضغط 📂 تحميل المشروع لاستعادة كل شيء كما تركته بالضبط.
هل مزامنة حركة الشفاه متوفرة؟
ليس بعد. مزامنة حركة الشفاه (مطابقة حركة الشفاه مع اللغة الجديدة) مخطط لها في تحديث مستقبلي. حالياً، ينتج التطبيق صوتاً مدبلجاً عالي الجودة يمكنك دمجه مع الفيديو الخاص بك.
هل ستتأثر دقة أو جودة الفيديو؟
لا. يتم نسخ مسار الفيديو الأصلي دون أي تغيير، ويُستبدل مسار الصوت فقط بالصوت العربي المدبلج (أو يُدمج معه). تبقى الدقة ومعدل البت ومعدل الإطارات وجودة الصورة كما هي في ملفك الأصلي.
هل يتم حفظ ملفاتي ومقاطع الفيديو المولدة على خوادمكم؟
لا. ملفاتك المرفوعة والنسخ الكتابية والصوت/الفيديو المُولّد تُخزّن مؤقتاً على الخادم فقط أثناء جلستك النشطة. تُحذف تلقائياً عند إعادة تشغيل الخادم أو بعد فترة قصيرة. يجب عليك تنزيل نتائجك أثناء جلستك — بمجرد إغلاق المتصفح أو إعادة تدوير الخادم، تختفي الملفات نهائياً. لا نحتفظ بنسخ من وسائطك. لسلامتك، اضغط دائماً على أزرار التنزيل في الخطوة ٦ قبل مغادرة الصفحة.