SmartSub is a desktop application for generating, translating, and dubbing video subtitles with local speech-to-text processing and subtitle burning.
The tool addresses the workflow of converting video audio to subtitles, translating them, applying text-to-speech dubbing with voice cloning, and burning subtitles into video files. It uses local models based on Whisper and other speech recognition engines to process audio offline, keeping files on the user's machine. The application supports batch processing and GPU acceleration across NVIDIA, AMD, Intel, and Apple Silicon hardware, running on Windows, macOS, and Linux. It also includes online video downloading, allowing users to paste links from platforms like YouTube or Bilibili and process them directly within the application.
The tool suits creators working with multilingual content, educators preparing subtitled lectures, and anyone processing podcasts or recordings into searchable subtitle files. The complete pipeline can run entirely offline and free of charge using local models, built-in translation sources, and local text-to-speech synthesis without requiring API keys or usage limits. For users needing additional capabilities, the tool optionally integrates with twenty translation services, nine cloud speech-to-text providers, and six cloud dubbing services. The application allows independent use of each step—downloading, transcription, translation, proofreading, dubbing, and export—or chaining them together for batch workflows.
Development activity shows consistent engagement with the codebase through regular updates and maintenance of the feature set. The project maintains active support across multiple platforms and continues to expand integration options with external services while preserving the core offline-first functionality.