When building Summarize — the 9th tool in our Craft suite — we needed to answer one question: which engine actually works well for both English and Vietnamese?
Instead of trusting paper scores, we tested with 4 real news articles. The blunt verdict: DeepSeek wins decisively, while ViT5 and mT5 fail badly on Vietnamese — one invented a government official who does not exist, the other returned broken tokens. This is not "potential needing improvement"; these are real failures. Following our honest benchmark principles, a failure gets called a failure.
5 engines tested
| Engine | Type | Size | Cost |
|---|---|---|---|
| DeepSeek V4 Flash | API (AI) | 0 (cloud) | $0.27/1K calls |
| DistilBART-6-6 | ONNX local | 284 MB | Free |
| ViT5-base (VietAI) | ONNX local | 900 MB | Free |
| mT5-small | ONNX local | 300 MB | Free |
| TextRank | Pure JS | 0 | Free |
Methodology
We used 4 real long-form articles, each 800-1,200 words: 2 English (the EU AI Act, fusion energy) and 2 Vietnamese (Vietnam's 2026 economy, Vietnam's semiconductor industry). Every engine processed the same input with no per-engine prompt tuning.
Speed was measured on our dev machine (Apple M-series, Vulkan via MoltenVK) with local engines on ONNX Runtime, averaged over 3 runs after warm-up (ERPFit internal benchmark). DeepSeek was timed end-to-end through its API, so it is network-bound. Quality was graded by hand, checking each output sentence against the source to catch hallucinations and broken tokens.
Why we rejected ROUGE
ROUGE measures n-gram overlap between an output and a reference summary. The problem: it cannot detect hallucination. The ViT5 output below scored around ROUGE-1 0.31 — sounds passable — yet it contains a fabricated official. A "mediocre" score hid a serious factual error. For an end user, a summary that states falsehoods is worse than no summary at all.
Speed
| Engine | English avg | Vietnamese avg | Notes |
|---|---|---|---|
| TextRank | 1ms | 1ms | Extractive, no AI |
| mT5-small | 166ms | 155ms | Fast but broken output |
| ViT5-base | — | 470ms | Vietnamese only |
| DistilBART-6-6 | 870ms | — | English only |
| DeepSeek V4 Flash | 17.6s | 7.2s | API, network-bound |
These speed figures are an ERPFit internal benchmark. Note: speed is meaningless if the output is wrong. mT5 is the fastest AI engine here, but it returns garbage.
Quality — actual outputs
Vietnamese: Vietnam Economy 2026
DeepSeek V4 Flash ✅ BEST
• GDP nửa đầu 2026 tăng 7,2%, vượt mục tiêu.
• Xuất khẩu đạt 210 tỷ USD; điện tử, bán dẫn đóng góp 45 tỷ USD (+28%).
• FDI đạt 18,5 tỷ USD; Samsung, Intel, NVIDIA mở rộng đầu tư.
• Lạm phát 4,5%; BĐS giảm 10-15%; VND mất 3%.
• Mục tiêu cả năm: tăng trưởng 7,5-8%.
ViT5-base ❌ HALLUCINATION (FABRICATED FACT)
0-7, Cục trưởng Cục Phát triển Nam Bộ ( Bộ Công thương ) Nguyễn Văn Đức cho biết, tăng trưởng kinh tế có những diễn biến tích cực...
❌ "Director Nguyễn Văn Đức" appears nowhere in the source. The model fabricated a named government official, with a title and an agency — all of it false. This is a hallucination, the most dangerous failure mode for a summarizer, because it produces convincing fake information. For Vietnamese, ViT5-base is not trustworthy for production use.
mT5-small ❌ BROKEN OUTPUT
<extra_id_0> lớn. /.. /... /......
❌ This is not a summary. mT5-small returned the sentinel token <extra_id_0> left over from pre-training, plus meaningless punctuation. The output is completely unusable — the model failed at Vietnamese summarization.
Vietnamese: Vietnam Semiconductors
The result repeated exactly: DeepSeek produced a coherent summary of the supply chain and FDI inflows into chips. ViT5-base again inserted phrasing not in the source, and mT5-small again returned a broken <extra_id_0> string. This is not a fluke — it is a systematic failure of both small models on Vietnamese.
English: AI Regulation 2026
DeepSeek V4 Flash ✅ BEST
• EU AI Act bans unacceptable-risk AI; high-risk needs conformity assessments; fines up to €35M / 7% turnover
• Compliance costs $50-100M per major tech company; startup consolidation wave
• Transatlantic divide: US voluntary, China own framework
• OECD proposes mutual recognition; early-stage negotiations
DistilBART-6-6 ✅ GOOD
The European Union's Artificial Intelligence Act took full effect in February 2026. The regulation establishes a risk-based framework that classifies AI applications into four tiers. Companies face fines of up to 35 million euros or 7 percent of global annual turnover.
mT5-small ❌ BROKEN OUTPUT
<extra_id_0> to the Artificial Intelligence Act. Artificial Intelligence.com/.
❌ Same failure as in Vietnamese: leaked sentinel tokens and fractured text. mT5-small is unusable for summarization in both languages.
English: Fusion Energy
DeepSeek and DistilBART both handled the fusion energy progress article well, capturing the commercialization timeline and investment figures. mT5-small again leaked tokens. TextRank extracted the right sentences but could not condense the ideas.
Quality rankings
| Engine | English | Vietnamese | Verdict |
|---|---|---|---|
| DeepSeek V4 Flash | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | Best for both languages |
| DistilBART-6-6 | ⭐⭐⭐⭐ | — (EN only) | Good English offline option |
| TextRank | ⭐⭐⭐ | ⭐⭐⭐ | Instant fallback, any language |
| ViT5-base | — (VI only) | ⭐ | Fabricates facts — untrustworthy |
| mT5-small | ⭐ | ⭐ | Broken token output — unusable |
Cost
DeepSeek V4 Flash costs only ~$10/year at 100 summaries/day, at $0.27/1K calls (ERPFit internal benchmark). Local engines are free, but for Vietnamese DeepSeek is not just better — it is the only option that produces trustworthy results. We cover how we keep AI costs low in our AI business integration guide.
Conclusion
Auto engine chain: DeepSeek → DistilBART → TextRank
- DeepSeek V4 Flash — default when online. Best quality for both EN+VI.
- DistilBART-6-6 — offline fallback for English, ~1 second.
- TextRank — instant fallback for any language when both unavailable.
We deliberately kept ViT5-base and mT5-small out of the automatic Vietnamese chain. ViT5 fabricates facts and mT5 returns broken tokens — both are too risky to use by default. Users can still pick them manually, but we will not hide the fact that they failed this test.
This same blunt approach drives our other benchmarks, such as the PDF compression benchmark.