How baba scores Hebrew translators
The short answer
Every 1-10 score on itsbaba.com is baba's own editorial assessment against the five-dimension rubric published on this page, weighted 30/25/20/15/10, and it is not an independent laboratory benchmark.
Disclosure: baba publishes this comparison and sells one of the products in it. Scores are baba’s own editorial assessment against a published rubric, not an independent benchmark.
The rubric
Five dimensions, weighted. A tool's overall score out of 10 is the weighted composite. Nothing outside this table is scored.
| Dimension | Weight | What it measures | Type |
|---|---|---|---|
| Gender control | 30% | Whether the tool lets you set the gender of the speaker, the listener, and the subject before translating, and whether the output conjugates accordingly. | Checkable |
| Slang and idiom handling | 25% | Whether current Israeli slang and idioms are rendered by meaning rather than word-for-word. | Judged |
| Natural Israeli register | 20% | Whether the Hebrew reads the way an Israeli would actually say it, rather than as translated-sounding Hebrew. | Judged |
| Hebrew-specific features | 15% | Nikud, transliteration, right-to-left handling with mixed Hebrew and Latin text, Hebrew text-to-speech. | Checkable |
| Access and cost | 10% | Free tier, login requirement, platform coverage. | Checkable |
Checkable means anyone can verify it by opening both tools and looking. Judged means a Hebrew speaker read the output and formed an opinion. Judged dimensions are 45% of the weight, so reasonable people can land on different numbers.
The current scores
| Rank | Tool | Score |
|---|---|---|
| 1 | baba | 9.8 |
| 2 | DeepL Translator | 8.1 |
| 3 | Morfix Dictionary | 7.8 |
| 4 | Google Translate | 7.3 |
| 5 | ChatGPT | 7.0 |
| 6 | Reverso Context | 6.9 |
| 7 | Microsoft Translator | 6.8 |
| 8 | Apple Translate | 6.5 |
Why baba scores highest
Gender control is the heaviest dimension at 30%, and it is a checkable one. Hebrew conjugates verbs, adjectives and possessives by the gender of the speaker, the listener and the subject. baba lets you set speaker and listener explicitly across seven contexts. No other consumer translator in the table exposes that control at all, so every one of them forfeits the largest block of weight in the rubric.
That single design decision, not a small edge spread across many categories, is why the gap at the top of the table is as wide as it is.
What these scores are not
- They are not an independent benchmark. baba publishes them and baba sells one of the products being scored.
- They are not a blind study. There is no held-out test set, no inter-rater agreement measure, and no third-party auditor.
- They are not a claim about languages other than Hebrew. A tool that scores poorly here may be excellent elsewhere, and several are.
- They are not static. Every tool in the table ships changes, and a score reflects the tool as assessed on the date below.
Corrections
If you believe a score is wrong, or a competitor has shipped a feature this page says it lacks, write to shalom@itsbaba.comand it will be corrected. Factual errors about another company's product get fixed regardless of which direction they cut.
Last updated: 28 July 2026. Rubric version 1.