wenyi
Carry stories across languages.
Translation for books and long-form writing, with the whole work in view.
Whole-book understanding · Consistent terminology · Evidence-based review
Quick start · Language support · Documentation
English | 简体中文
- Why Wenyi
- Core features
- Interface preview
- Quick start
- Supported formats
- Translation pipeline
- Documentation
- Limitations
- Community
- Support
- Star history
- License
Why Wenyi
| Typical approach | Wenyi |
|---|---|
| Segments translated in isolation, unaware of surrounding content | Whole-book prescan with chapter digests and rolling context |
| Glossary managed manually or as an afterthought | Real-time term extraction with conflict detection, fed back into subsequent batches |
| Single-pass translation, fragile to interruptions | Batch checkpoints and chapter status tracking: resume any interrupted run with the same command |
| Raw model output, no systematic quality process | Translate → polish → evidence-driven whole-book review |
Wenyi is designed for long-form texts — novels, social-science monographs, narrative nonfiction, and more.
A bilingual reading sample: translation alongside visually subdued source text.
Core features
- Web workspace — English and Chinese interfaces, live translation progress, paragraph proofreading with revision history, and whole-book review with evidence and publication results.
- Whole-book understanding — prescans the source before translation, creating per-chapter digests and a book-level synopsis injected into every batch
- Real-time glossary — extracts proper names, terms, and recurring expressions as translation progresses; detects conflicting translations and surfaces them for resolution
- Multi-stage quality — optional polishing (strong model) and an evidence-driven whole-book AI review
- Resumability — batch-level checkpoints, chapter status tracking, and atomic state writes; interrupt at any point and resume with the same command
- Multiple LLM providers — DeepSeek, OpenAI, OpenRouter, OrcaRouter, Google Gemini, Ollama, vLLM, and generic OpenAI-compatible endpoints; keep three convenient tiers or select models per operation, mix connections, and share request limits. See model routing.
- Native EPUB preservation — writes translated text back into the original XHTML templates and attempts to preserve styles, images, TOC, and anchors
- Bilingual output — optional source-and-translation edition with visually subdued source text, including dark mode support
Interface preview
Track translation progress, usage, and elapsed time, then proofread paragraphs alongside the source. See the deployment guide. Screenshots show the Chinese interface; English is available in Settings.
Translation overview: usage by step, cache hit rates, and run durations.
Manual proofreading: compare source and translation; right-click to edit, inspect revisions, or copy text.
Quick start
Prerequisites
Wenyi requires Python 3.10+ and uv.
Installation
git clone https://github.com/BigDawnGhost/wenyi.git
cd wenyi
uv sync
Configuration
Set your API key:
export DEEPSEEK_API_KEY=sk-...
One-command translation
uv run wenyi translate book.epub
This parses the book, detects the source language, prescans for understanding, translates all chapters, and assembles the output. The monolingual Chinese EPUB is written to output/book.zh.epub by default.
Multilingual translation (experimental): select a direction using language.source / language.target, such as zh → en or en → ja. Run uv run wenyi languages for the list. Targets have separate state and output names. See the usage guide.
Step-by-step workflow
# 1. Prepare — parse, analyze, prescan (no body text translated)
uv run wenyi prepare book.epub
# 2. Translate — resume from the prepared state
uv run wenyi translate book.epub
# 3. Review — independent final review against the completed glossary
uv run wenyi review book.epub
# 4. Check progress
uv run wenyi status book.epub
Interrupt and resume
Every completed batch is persisted immediately. If a run is interrupted, execute the same command again:
uv run wenyi translate book.epub
Command-line overrides
uv run wenyi translate book.epub --polish --review # enable polishing and final review
uv run wenyi translate book.epub --no-polish # disable polishing
uv run wenyi translate book.epub --no-review # skip final review
uv run wenyi translate book.epub --bilingual # produce both editions
uv run wenyi translate book.epub --chapter 0 # translate the first chapter (indices start at 0)
uv run wenyi translate book.epub --format txt # export as plain text
Final review runs by default after the complete book has been translated and the
glossary has reached its final state. Pass --no-review or set
pipeline.review: false to skip it. You can also run Agent Review independently:
uv run wenyi review book.epub
uv run wenyi review book.epub --autofix
Translation batches and fresh Reviewer requests receive the full glossary. Review reuses completed results or resumes unfinished work when content, configuration, and full-glossary fingerprints match; otherwise it starts a new run. It checks chunks concurrently and can selectively request cross-book evidence before resolving contradictory consistency suggestions. See glossary and resume policy.
Confirmed issues can produce provisional full-segment
replacements in a run-local shadow translation. A fresh whole-book review sees
the shadow text—but not the previous issue explanation—and validates it again.
Review publishes to formal chapter target values by default. Pass
--no-autofix or set pipeline.review_autofix: false to keep the run
read-only. With Autofix, folded changes are applied first and remaining
issues reuse the existing Review Agent Loop and Fixer against that updated text.
Only formal segment target values are replaced; full history stays in the Review
directory's autofix/index.json. The consolidated result, run usage, events, and
internal records are written under state/<book>/targets/<target-language>/reviews/review-<timestamp>/.
Supported formats
| Input | Output |
|---|---|
| EPUB, FB2, TXT, Markdown, HTML, PDF, DOCX | EPUB (monolingual / bilingual), TXT, HTML, Markdown, DOCX |
| SRT (movie / series subtitles) | .zh.srt (monolingual) and optional .zh-bi.srt (bilingual) |
- PDF input defaults to MinerU and requires
MINERU_API_KEYfor the initial conversion; the resulting HTML is cached and reused. The BabelDOC bridge is optional for layout-preserving PDFs. - EPUB output attempts to preserve the original book's styles, images, table of contents, and anchors. Vertical layout is converted to horizontal for Chinese reading.
- Source language is auto-detected by default, or fixed to an ISO 639-1 code in
config.yaml. .srtinput is auto-detected bytranslate. It uses a light concurrent path (no glossary, polish, or whole-book review). State lives understate/srt/<slug>/targets/<target-language>/; outputs default to the source file'soutput/directory. Details: Usage guide..docxinput uses the full book pipeline. Headings, simple tables, lists, and common run/paragraph styles are preserved where possible; translated Chinese uses Song (宋体). Default export is.zh.docx(override with--format). Details: Usage guide.
Translation pipeline
Wenyi combines whole-book understanding, batch translation, optional polishing, review, and export. See the translation pipeline for the flowchart and stage details.
Documentation
- Usage guide — installation, Windows setup, input/output, resumability, independent stages
- Configuration — providers, languages, pipeline switches, segmentation, paths
- Translation pipeline — whole-book analysis, terminology, context, polishing, review
- Web deployment — Docker/local Web stack, workers, exports, and project workflows
- Contributing — development, testing, and contribution guidelines
Translated state directories for public-domain books may be shared through wenyi-bookcase. Do not publish copyrighted text, private books, or state/ directories containing sensitive information without permission.
Limitations
- Multilingual translation is experimental: Chinese, English, Japanese, Korean, French, German, Spanish, Italian, Portuguese, Russian, Vietnamese, and selected variants have built-in profiles. Real-model long-form quality still needs evaluation; the CLI and prompt instructions use English, while generated descriptive metadata follows the translation target.
- Polishing and final review are the most expensive stages. Shadow fixing may trigger multiple full-book review passes and additional Fixer calls.
- PDF input defaults to MinerU and requires an API key for the initial conversion. The BabelDOC bridge is optional for layout-preserving PDFs.
- SRT translation is a light concurrent path: no glossary, polishing, or whole-book review, and slug collision is possible for identically named files in different folders.
- Translation quality is bounded by the capabilities of the chosen LLM model.
- Very long books may produce large state directories; storage requirements grow with book length.
Community
- Discord server
- QQ group: 1055065098
- GitHub Issues — bug reports and feature requests
- GitHub Discussions — ideas and questions
Support
If this project has been helpful, tips are welcome.
WeChat Pay · Alipay
Star history
License
AtomGit (China)
Wenyi is also hosted on AtomGit: https://atomgit.com/BigDawnGhost/wenyi
