Homework built from your own mistakes
The problem with a workbook
A printed workbook has to work for everyone who buys it, which means it is calibrated to a hypothetical average learner who does not exist. You will do exercises on structures you already handle fine, and you will do them at the same length and difficulty as the structures that are genuinely hard for you. Some of that is useful review. Most of it is time spent proving something you had already proved.
The same problem shows up in an app that serves everyone the same exercise queue. The exercises may be well made. They are still aimed at nobody in particular.
The only source of exercises genuinely about you is the record of what you got wrong while speaking — not a placement test, and not a diagnostic quiz taken once at the start.
Where the material comes from
In a lesson, you produce sentences under time pressure, and some of them are wrong. Those errors are the raw material. After the session, the analysis step files each one under a type — articles, tenses, prepositions, word order — quotes the sentence you actually said, and notes how often the pattern recurs. Separately it logs the vocabulary you reached for and did not have, which is a different and equally useful kind of gap: not something you got wrong, but something you could not get to.
Homework is generated from that record rather than from a shared pool. In practice that means a handful of exercise kinds, each fed by a different part of it.
Flashcards come from the vocabulary queue — words and phrases you missed or newly met, scheduled to come back at widening intervals. Grading here is deterministic: a word is recalled or it is not, and no model is needed to judge that.
Cloze and transformation exercises are generated against your specific error patterns. If your record says articles, you get sentences with the articles removed and sentences to rewrite. If it says present perfect against past simple, you get the pair contrasted in contexts where the choice actually matters. Grading is mostly deterministic with some tolerance for equivalent answers, because “I have lived here since 2020” and “I have been living here since 2020” are both defensible and a rigid string match would fail one of them.
Listening clips with questions, generated as audio and themed on your interests. Listening is not the bottleneck this product is built around, but comprehension and production feed each other, and a clip about a subject you care about gets finished.
Short readings with questions, also themed. Same reasoning.
Free writing, a few sentences against a lesson objective. Writing is slower than speaking and therefore easier to be accurate in, which makes it a decent way to consolidate a structure you can produce carefully but not yet fluently.
Pronunciation drills, where you record yourself and the result is compared against a target. This is the least settled of the set — automatic pronunciation scoring is genuinely hard and over-confident scoring is worse than none — so it belongs in the category of things to treat as experimental rather than as a grade.
Timing is most of the value
Getting the content right matters less than getting the timing right, which is the part most homework systems ignore entirely.
An exercise given immediately after the lesson is nearly free. The material is still in working memory, you get it right, and almost nothing is consolidated — you have practiced remembering something you had not yet forgotten. An exercise given a month later, when the item is gone, is not review either; it is re-learning from scratch, at full cost.
The useful moment is in between: the point where recall is effortful but still possible. So each item carries its own schedule rather than sitting in one global queue. Get it right and the interval widens — roughly a couple of days, then closer to a week, then a couple of weeks, then a month. Miss it and the interval resets short, because a miss is evidence the previous interval was too long for that item.
This is per item, not per learner and not per topic. Two vocabulary items met in the same lesson can end up on completely different schedules within a fortnight, because one of them stuck and one did not. A single queue with a single interval cannot express that, and the difference compounds over a course.
The part that closes the circle
Homework that produces a score and stops is still an open loop. What makes it part of a course is that the results go back into the same record they came from.
A completed exercise updates two things. It updates the review schedule for the items it touched — recalled or missed, next due date set accordingly. And it updates the mastery state of the curriculum objective behind it, which is what the next lesson plan reads when it decides difficulty.
Then a summary of what you did goes into the next conversation. Not as a report read out to you, but as context the tutor has before it opens its mouth: this learner did the article exercises and got most of them, and still missed the definite article with abstract nouns. That is a lesson that can start somewhere specific.
The reverse direction matters too. A structure you drilled correctly on a screen is not a structure you can produce out loud under time pressure — those are the two different skills that make speaking hard in the first place. So homework is not the end of the item’s life. It is the rehearsal before the item gets tested where it counts, in the conversation, where you have no time to think about which tense you are choosing.
Why it is not announced as a test
There is a strong temptation to make the review visible: “let us review the mistakes from last time.” It is honest and it is a mistake. Announcing a drill changes what the learner is doing — they switch from speaking to performing, they slow down, and they produce the careful version of their English rather than the real one. You end up measuring their best effort under exam conditions, which is not the thing you were trying to improve.
The better approach is to steer the conversation toward a situation that naturally requires the structure, and see what comes out. If the item is the past simple for onset of symptoms, a doctor role-play will demand it without anyone mentioning grammar. If it is making requests and asking for information, a hotel check-in will. The scenario library exists partly for this: each brief carries the structures it naturally exercises, so a review item can be matched to a situation rather than to a worksheet.
You still find out whether you have it. You just find out under conditions that resemble the ones you actually care about.
What good homework feels like
Short, specific, and slightly annoying. Short because between-lesson work competes with everything else in your week and a long queue simply does not get done. Specific because every item should be traceable to something you personally got wrong or reached for and missed. And slightly annoying because if it is comfortable, the interval was too short and you are practicing recall you already had.
What it should never feel like is a random assortment. If you cannot see why a given exercise was assigned to you, either the system does not know, or it knows and did not bother to use it — and from the outside those look identical.
The closed learning loop post covers the full cycle this sits inside, and how it works walks through where homework lands between one lesson and the next.