← Back to App / العودة للتطبيق

Help & FAQ

المساعدة والأسئلة الشائعة

How It Works — Step by Step

Step 1: Upload

Upload an audio or video file (MP3, WAV, MP4, AVI, MKV, MOV, WEBM). Maximum 60 seconds and 400 MB. The app transcribes the English speech and identifies each speaker automatically.

Step 2: Edit Segments

Review the transcription. You can:

  • Edit text — fix any transcription mistakes directly in the table.
  • Auto Translate — translate all unlocked lines to Arabic using AI.
  • Add Tashkeel — add Arabic diacritics (harakat) to the translated text.
  • Detect Emotions — AI listens to each segment's voice tone and assigns a speaking style.
  • Auto-Fix Timing — re-sync segment timestamps to match the original audio precisely.
  • 🔒 Lock — protect specific lines from being changed by Auto-Fix, Translate, or Tashkeel.

Step 3: Voice Cloning Setup

The app analyzes how much speech each speaker has. Speakers with enough audio (≥1 second) can be cloned — the AI copies their unique voice from the video so the Arabic dub sounds like them.

Step 3.5: Choose Speakers to Clone

Review the quality guidance for each speaker. Uncheck any speaker you don't want to clone. Speakers with very little audio will produce poor clones — use a studio library voice instead.

Step 4: Assign Voices

Pick a voice for each speaker. You can use:

  • Cloned voices — copied from your video (if cloning succeeded).
  • Studio library voices — high-quality professional voices (numbered, names hidden).

Use 🎲 Auto-Assign to let the app pick the best match automatically. Two speakers never share the same numbered voice.

Step 5: Generate Arabic Audio

Click Generate Arabic Audio. The AI speaks each line in Arabic using the assigned voice and emotion/style. This may take 1–3 minutes depending on the clip length.

Step 6: Review & Fine-Tune

Listen to the result. You can:

  • 🔄 Re-speak a single line — regenerate just one line without redoing the whole clip.
  • 🎚️ Fine-Tune Timeline — drag segment blocks left/right (in milliseconds) to perfectly align with lip movements. Blocks cannot overlap.
  • 🎬 Merge into Video — combine the dubbed audio with the original video and background music.
  • 💾 Save / Load Project — save your work at any time as a JSON file. Load it later to continue editing without re-transcribing.

Frequently Asked Questions

What file formats are supported?

Audio: MP3, WAV, FLAC, AAC, OGG. Video: MP4, AVI, MKV, MOV, WEBM. Maximum duration: 60 seconds. Maximum file size: 400 MB.

How does voice cloning work?

The AI extracts each speaker's voice characteristics from the original audio and creates a synthetic copy. This copy is then used to speak the Arabic translation in a voice that sounds similar to the original speaker. At least 1 second of clear speech is required; 10+ seconds produces the best results.

What are credits and how are costs calculated?

Credits are our internal billing unit: 1 credit = $0.01 USD (100 credits = $1.00). Translation and analysis cost a few cents per clip. Voice generation is charged by character count. Every button shows its estimated cost before you click it, and actual usage appears in the green box in Step 1.

What does the 🔒 lock do?

Locking a line protects it from three actions: Auto-Fix Timing, Auto Translate, and Add Tashkeel. Locked lines keep their current text, timing, and formatting untouched. Use this for lines you've manually perfected.

Can I combine multiple emotion/style tags?

Yes! Use the dropdown to add tags one at a time, or type directly in the text box (e.g., "confident, calm"). Only recognized tags from the official list are accepted; unrecognized words are removed automatically. Examples: "anxious, afraid", "playful, teasing", "calm, firm".

Why does my cloned voice sound robotic?

This usually means the speaker didn't have enough clear audio in the original clip. Check the quality guidance in Step 3.5 — if it says ⚠️ or ❌, switch that speaker to a studio library voice in Step 4 for better results.

How does the timeline fine-tuning work?

In Step 6, click 🎚️ Fine-Tune Timeline. Each blue block represents one spoken line. Drag blocks left or right to shift their timing by milliseconds. A time ruler at the top shows your position. Blocks stop at adjacent segments to prevent overlap. Click ✅ Confirm to rebuild the audio with your adjustments.

Does the merged video keep the original background music?

Yes. The app separates vocals from background audio during processing. When you merge, the dubbed Arabic vocals are mixed with the original background music/sound effects at balanced volumes.

Can I save my progress and come back later?

Yes. Click 💾 Save Project in Step 2 to download a JSON file with all your segments, translations, voice assignments, and settings. Later, click 📂 Load Project to restore everything exactly where you left off.

Is lip-sync available?

Not yet. Lip-sync (matching lip movements to the new language) is planned for a future update. Currently, the app produces high-quality dubbed audio that you can merge with your video.

Will the resolution or quality of my video be affected?

No. The original video stream is copied unchanged; only the audio track is replaced (or mixed) with the dubbed Arabic audio. Resolution, bitrate, frame rate, and visual quality stay exactly as your source file.

Are my uploaded files and generated videos saved on your servers?

No. Your uploaded files, transcriptions, and generated audio/video are stored temporarily on the server only while your session is active. They are automatically deleted when the server restarts or after a short period. You must download your results during your session — once you close the browser or the server recycles, the files are gone permanently. We do not keep copies of your media. For your safety, always click the download buttons in Step 6 before leaving the page.