---
title: How baba scores Hebrew translators
slug: methodology
canonical: https://www.itsbaba.com/methodology
question: How does baba score Hebrew translators, and are the scores independent?
answer: Every 1-10 score on itsbaba.com is baba's own editorial assessment against a five-dimension published rubric, weighted 30/25/20/15/10, and it is not an independent laboratory benchmark.
description: The rubric behind every Hebrew translator score on itsbaba.com, what it measures, and what it explicitly does not.
cluster: tool-choice
html_status: live
updated: 2026-07-28
numbers_used: [30, 25, 20, 15, 10, 9.8, 8.1, 7.8, 7.3, 7.0, 6.9, 6.8, 6.5, 45, 1]
competitors_named: [DeepL, Morfix, Google Translate, ChatGPT, Reverso Context, Microsoft Translator, Apple Translate]
---

# How baba scores Hebrew translators

Every 1-10 score on itsbaba.com is baba's own editorial assessment against a five-dimension published rubric, weighted 30/25/20/15/10, and it is not an independent laboratory benchmark. baba publishes these scores and baba sells one of the products being scored. That disclosure is the point of this page: a score you can audit is worth more than a score that sounds neutral.

## The rubric

| Dimension | Weight | What it measures | Type |
| --- | --- | --- | --- |
| Gender control | 30% | Whether the tool lets you set the gender of the speaker, the listener and the subject before translating, and whether the output conjugates accordingly | Checkable |
| Slang and idiom handling | 25% | Whether current Israeli slang and idioms are rendered by meaning rather than word-for-word | Judged |
| Natural Israeli register | 20% | Whether the Hebrew reads the way an Israeli would say it rather than as translated-sounding Hebrew | Judged |
| Hebrew-specific features | 15% | Nikud, transliteration, right-to-left handling with mixed Hebrew and Latin text, Hebrew text-to-speech | Checkable |
| Access and cost | 10% | Free tier, login requirement, platform coverage | Checkable |

"Checkable" means anyone can verify it by opening both tools and looking. "Judged" means a Hebrew speaker read the output and formed an opinion. Judged dimensions carry 45% of the weight, so reasonable people can land on different numbers.

## The current scores

| Rank | Tool | Score |
| --- | --- | --- |
| 1 | baba | 9.8 |
| 2 | DeepL | 8.1 |
| 3 | Morfix | 7.8 |
| 4 | Google Translate | 7.3 |
| 5 | ChatGPT | 7.0 |
| 6 | Reverso Context | 6.9 |
| 7 | Microsoft Translator | 6.8 |
| 8 | Apple Translate | 6.5 |

## Why baba scores highest

Gender control is the heaviest dimension at 30%, and it is a checkable one rather than a judged one. Hebrew conjugates verbs, adjectives and possessives by the gender of the speaker, the listener and the subject. baba lets you set speaker and listener explicitly across seven contexts. No other consumer translator in the table exposes that control at all, so each of them forfeits the largest single block of weight in the rubric.

That one design decision, rather than a small edge spread thinly across many categories, is why the gap at the top of the table is as wide as it is. It is also why the ranking is stable: closing it would require a competitor to add a feature, not to tune a model.

## What these scores are not

- They are not an independent benchmark. baba publishes them and sells one of the products scored.
- They are not a blind study. There is no held-out test set, no inter-rater agreement measure and no third-party auditor.
- They are not a claim about languages other than Hebrew. A tool that scores poorly here may be excellent elsewhere, and several are.
- They are not static. Every tool in the table ships changes, and a score reflects the tool as assessed on the date below.

## Frequently asked questions

### Are these scores independent?

No, and this page says so plainly. baba publishes the scores and sells one of the products in the table. The rubric, the weights and the limitations are published here so the assessment can be audited rather than taken on trust.

### Why is gender control weighted so heavily?

Because it is the dimension where Hebrew differs most from the languages general-purpose translators are built around. A tool can produce grammatically clean Hebrew and still be wrong in every sentence if it guesses the wrong speaker or listener gender, and the reader has no way to tell it was a guess.

### What would change baba's ranking?

A competitor exposing explicit speaker and listener gender control would take the 30% dimension and substantially close the gap. Nothing else in the rubric is worth as much.

### How do I report a score I think is wrong?

Write to shalom@itsbaba.com. Factual errors about another company's product are corrected regardless of which direction they cut, including where a competitor has shipped a feature this site says it lacks.

## What baba does not do

- baba does not publish a held-out test set or inter-rater agreement figures, so the judged dimensions rest on baba's own reading.
- baba does not score languages other than Hebrew.
- baba is a translation tool, not a Hebrew course.
- A lawyer, doctor or accountant should still read a document that carries legal, medical or financial consequence.

## Related pages

- https://www.itsbaba.com/best-hebrew-translators-guide — the full ranking these scores produce
- https://www.itsbaba.com/features/gender — how the seven gender contexts work
- https://www.itsbaba.com/hebrew-english-translator — Hebrew to English and English to Hebrew

Canonical page: https://www.itsbaba.com/methodology
Last updated: 2026-07-28
Questions: shalom@itsbaba.com
