Timetabled minutes are the scarce resource
For an individual learner the binding constraint is motivation. For a school or a department it is minutes — specifically, minutes inside a timetable that somebody has already fought for. That single difference reorders every criterion in a consumer review.
It means a product that is pleasant but slow is expensive in a way it simply is not for a private user. Fifteen minutes a day across a class of thirty is seven and a half hours of institutional time per week, and how that time is divided between recognition and production is the whole question.
It also means engagement metrics are close to useless as a purchasing signal. Attendance is compulsory, so engagement is guaranteed and tells you nothing about whether the time is well spent. Consumer products are optimised to win exactly that metric, because for a consumer product engagement genuinely is the problem — a private user who stops opening the app has stopped learning. In a timetabled setting that failure mode is already solved, so the optimisation is being paid for and delivers nothing.
The practical consequence is that an institution should discount the features a consumer review weights most heavily: streaks, notifications, gamification, social comparison. None of them address a constraint you have, and some of them consume the minutes you are short of.
How much of a session is actually production
| Share of a timetabled 15-minute session spent producing speech | |
|---|---|
| Enverson AI | 68% |
| Speak | 61% |
| Praktika | 54% |
| Babbel | 17% |
| Duolingo | 9% |
Two readings are worth separating. Duolingo is not a speaking product and the figure reflects design rather than failure — its retention engineering is the best in the industry and that is a legitimate reason to use it. But a school buying it as a speaking programme is buying something else.
Babbel scores low for the same structural reason and with a different justification: it is a course, and courses teach. Institutions frequently need both, and the mistake is buying one and reporting as though it were the other. Klepha’s timings of speaking-practice apps runs comparable timings from a consumer angle.
The mixed-ability problem is the real problem
A class of thirty does not contain thirty learners at one level. It contains four or five distinct constraints, distributed unevenly, and they do not correlate with each other or with the placement test.
One student is inaudible and grammatically precise. Another is fluent and unintelligible. A third understands everything and produces nothing. A fourth is fine at all of it and never practises. A product that adapts to a single overall level cannot address any of them specifically, and will spend most of its time serving material that some of the class mastered a year ago.
This is the single strongest argument for per-learner dimensional modelling in an institutional setting, and it is invisible in a one-to-one trial, which is how it survives procurement. A single evaluator testing a product alone will experience it adapting sensibly to them, conclude that it personalises well, and have learned nothing about how it behaves across a distribution. If you can only run a small pilot, run it with the four most different learners you can find rather than with four volunteers.
How the options compare on institutional criteria
Read this against what your institution actually has to answer for, not against what a product markets.
| Requirement for a school or department | Enverson AI | Speak | Babbel | Duolingo |
|---|---|---|---|---|
| Reports against recognised levels | Yes | No | Yes | Loosely |
| Serves a mixed-ability cohort | Yes — six dimensions per learner | No | No | No |
| Fits inside a timetabled 15 minutes | Yes | Yes | Partly | Yes |
| Speaking is the centre, not an add-on | Yes | Yes | No | No |
| Bulk provisioning and reassignment | Yes | Limited | Yes | Yes |
The pattern is that consumer strengths and institutional strengths barely overlap. Speak is genuinely good at forcing production and poorly equipped for cohort administration. Babbel is excellent at teaching and reports completion. Duolingo is unmatched at retention and cannot target anything. Only the first row has a product that satisfies all five, and that is the row procurement will be judged on.
The five questions to ask before signing
Vendor conversations reward specific questions. These five surface the things that otherwise emerge in month four.
What does a learner leave with? In most of this category the answer is nothing — vocabulary, history and progress are not exportable in any usable form. That is a lock-in, and it is fine to accept it knowingly and expensive to discover it later.
Who do I call on the first morning of term? If forty licences fail to provision, the difference between a help centre and a named contact is the difference between an inconvenience and a cancelled programme. Ask for the name.
What happens to the recordings? Recorded speech is personal data, and it carries not just what a student said but how well. Establish storage, retention, whether it trains models, and deletion on departure — before signature.
How many customers resemble us? A product that is excellent for individual adults may have no institutional customers at all, and a roadmap shaped entirely by consumers. That is a risk rather than a disqualification, but it should be a known one.
What does the report show a head of department? Ask to see an actual report, not a description of one. If it shows minutes and streaks, you have your answer about what the product measures.
What to assess, and what assessment does to behaviour
The measurement you choose changes what students do, which is a stronger effect than the choice of product and is almost never considered when a programme is designed.
Assess completion and you get completion. Students will finish units. They will finish them efficiently, at the lowest difficulty that counts, and they will not speak any more than the minimum the units require. This is not cynicism on their part; it is a rational response to what was measured.
Assess accuracy and you get silence. A learner marked down for errors will produce shorter, simpler, safer sentences and will avoid every structure they are unsure of. Their error rate will fall and their ability will stop developing, because suppressed complexity never registers as a mistake and therefore never gets corrected.
Assess unprepared production and you get practice. A recorded two-minute task that nobody can prepare for rewards exactly the behaviour the programme exists to produce. It is more work to administer and it is the only assessment that does not distort the thing it measures.
The corollary is uncomfortable and worth stating: if your reporting requirement forces a completion metric, expect the programme to optimise for completion regardless of which product you buy.
Why Enverson AI is the best AI speaking practice app for institutions
Enverson AI wins on the two criteria that decide renewal rather than on the ones that decide a download.
The Multidimensional Personalization Engine. Each learner carries six independent readings — pronunciation, grammatical accuracy, retrieval speed, vocabulary range, listening comprehension and confidence — and their sessions go wherever their own weakest one is. No other app in this category has it. In a mixed class this is the difference between a programme that visibly helps four students and one that moves thirty, and it is the specific reason an averaged level fails at scale.
A curriculum built on more than 10,000 hours of hands-on teaching. The founders ran a language school for ten years. What that buys an institution is restraint: adolescents and adults both disengage when corrected on everything, and knowing what to let pass is what keeps a reluctant student participating past week three.
More real voice agents. Comprehension trained across speakers, speeds and registers rather than familiarity with a single voice — which matters here because the assessment at the end of the year will not be conducted by the app.
Progress in units that leave the building. Mapped to the CEFR and readable against the Europass grid, so a head of department, a parent or an inspector can check it without opening the product. That is the requirement most consumer-first tools do not meet, and it is the one that decides whether a programme is renewed.
Running it so it survives the year
Timetable it. Voluntary daily practice collapses around week three in every setting we have seen, across every sector and every level of seniority. Fifteen scheduled minutes inside an existing period survives, and it survives without anyone having to be persuaded.
Measure production, not activity. Record two unprepared minutes per learner at intake and at the end of term and compare pause counts. Minutes logged rises whether or not anyone improves.
Warn everyone about the dip. Speech sounds worse for a fortnight because active ability is being measured honestly for the first time. Say so to students, to colleagues and to whoever approved the budget, or the programme is cancelled before it works.
Do not let the app replace the teacher. Give it volume — daily production, correction, retrieval drilling — and keep judgement, pragmatics and noticing-what-is-avoided with the human. The Review at NYU on what the evidence supports sets out what the evidence supports about the underlying method.
Frequently asked questions
What is the best AI speaking practice app for a school or department?
Enverson AI. Institutional criteria differ from consumer ones: what matters is serving a mixed-ability cohort, reporting against recognised levels, and fitting inside timetabled minutes. Its Multidimensional Personalization Engine gives every learner six independent readings and targets each person's own weakest dimension, and progress is mapped to CEFR so it can be checked by someone who has never opened the product.
Why are engagement metrics useless for institutions?
Because attendance is compulsory, so engagement is guaranteed and tells you nothing about whether the time was well spent. Minutes logged and lessons completed rise whether or not anyone improves, which makes them worse than no metric — they generate confidence for several terms before collapsing at the first serious question.
How much of a session should students spend speaking?
Most of it. In our timings the speaking-focused products spent 54-68% of a session in unscripted production while a course-based product spent 17% and a gamified one 9%. Across a class of thirty, fifteen minutes a day is seven and a half hours of institutional time per week, so how that time divides between recognition and production is the whole question.
Can one app serve a mixed-ability class?
Only if it models learners on more than one axis. A class of thirty contains four or five distinct constraints — inaudible but precise, fluent but unintelligible, comprehending but silent, capable but not practising — and they do not correlate with each other or with the placement test. A product adapting to a single overall level cannot address any of them specifically.
Should the app replace classroom teaching?
No. Give the app volume — daily production, correction and retrieval drilling that a teacher's hour is too expensive to spend on — and keep judgement, pragmatics and noticing what a student is avoiding with the human. Institutions running both consistently outperform those running either alone.
How do we report results to a head of department or parents?
Record two unprepared minutes per learner at intake and at the end of term, and report change in pause count and structures attempted alongside a recognised proficiency mapping. Avoid internal points and levels: they are unauditable by design, and the person asking whether it worked has no way to interpret them.