Spend an hour in any GRE forum and the same five questions come round again and again. How accurate is PowerPrep? Why did I score lower on the real thing than in every mock I took? Are the Manhattan and Kaplan tests just harder on purpose? What score should I actually expect on test day? And the one nobody wants to ask out loud: is my target still realistic?
They are all really one question wearing five hats: what does a practice score actually tell me about my real score? This guide answers that properly. Not with reassurance, but with the mechanics of how these tests are scored, the specific reasons practice numbers inflate, and what to do with the answer once you have it.
One thing to settle first: a practice score is an estimate with an error bar, not a prediction. Treating it as a promise is the single most common way people set themselves up for a bad test day. If you want the mechanics of how the test adapts, read what section-adaptive actually means alongside this.
How accurate are PowerPrep scores?
This is the most repeated question in GRE communities, and it has a genuinely good answer: PowerPrep is the most accurate practice score you can get, at any price, and it is still routinely misused in ways that inflate it.
PowerPrep is written by ETS, the organisation that writes the real GRE. The questions are retired official items, the interface is the real interface, and the scoring runs the real section-adaptive algorithm. Nothing from a third party matches that. If you take it under honest conditions, a PowerPrep score is close to your real score more often than not. Our GRE PowerPrep guide covers what each test includes and how to get them free.
The catch is that PowerPrep only behaves like the real test if you make it behave like the real test. Here is what quietly breaks it:
| What people do | What it does to the score | The fix |
|---|---|---|
| Take PowerPrep 1 and treat the result as a baseline | There is no Verbal or Quant score to treat as anything. PowerPrep 1 is untimed, and ETS does not report V or Q scores for an untimed sitting. | Use PowerPrep 2, the timed one, for any score you intend to trust. |
| Pause the timer, take a phone call, stretch a break | Inflates the score, often by several points per section. Time pressure is a large part of what the GRE measures. | One sitting, one timer, only the scheduled break. |
| Split the test across two evenings | Removes the fatigue factor entirely. Section four is much harder when you have already been working for 90 minutes. | Full length, start to finish, in one block. |
| Retake a PowerPrep test they have already seen | Badly inflates it. You are partly measuring recall of the answer key, not ability. This is the biggest single source of fake scores. | Each official test is a one-shot instrument. Never re-sit one for a number you intend to believe. |
| Take it at 10pm on a laptop in bed | Mild inflation from comfort, and no rehearsal of the real conditions. | Match the real time of day and sit at a desk. |
If you have already used both official tests casually, you cannot buy another honest baseline from ETS for free. That is the practical reason to protect PowerPrep 2 until you actually want the number.
So the honest summary: a first-exposure, timed, uninterrupted, full-length PowerPrep 2 is a strong predictor. Anything else is a practice activity, and a practice activity is not a forecast.
Why did I score lower on test day than in my practice tests?
This is the most painful thread on any GRE forum, and it almost always has a findable cause. When someone posts that they mocked 325 and scored 312, the gap is rarely mysterious. It is usually one or more of the following, in rough order of how often they turn out to be the culprit.
- **The practice conditions were easier than test conditions.** Untimed sections, extra breaks, a familiar desk, no strangers, no check-in stress, and the quiet knowledge that the score does not count. Every one of those is worth points. Together they can be worth a lot.
- **Repeat exposure to the same questions.** If a mock was taken twice, or the questions overlapped with a book you had worked through, the score measured memory as much as reasoning. This is by far the most common cause of a large gap.
- **The scoring conversion was somebody's estimate.** Third-party tests do not have access to the real scaling. Their raw-to-scaled conversions are approximations, and several run generous. More on this in the next section.
- **Adrenaline and pacing.** A first section that feels harder than expected makes people rush, and rushing on the GRE produces careless errors on questions they can genuinely do. This is a real effect and it is trainable.
- **Fatigue that was never rehearsed.** If every practice session was a single 20-minute section, the fourth scored section on test day is a new experience. Stamina has to be practised at full length.
- **Misreading the adaptive design as failure.** Doing well on the first section of a measure routes you into a harder second section. A harder second section is a good sign, not a bad one, but it feels like collapse and people panic and lose points to the panic.
- **Ordinary score variance.** A handful of questions can swing a section score by a few points. Two honest attempts by the same person on the same day would not produce identical scores. A 2 to 3 point difference per section is noise, not a story.
Worth knowing if you are working from older advice: the current shorter GRE removed the unscored experimental section. The whole test now runs about 1 hour 58 minutes. Guides that warn you to budget stamina for a surprise extra section are out of date, though stamina still matters across the four scored sections.
Diagnosing which of these applies to you is the entire job, because the fix is completely different in each case. A gap caused by repeat exposure means your true level was always lower and you need more study. A gap caused by pacing means your level is fine and you need test-day execution work. Treating the second as the first wastes months.
Are Manhattan Prep, Kaplan and Princeton Review mocks accurate?
They are useful. They are not predictive in the way people want them to be, and the reason is structural rather than a criticism of the companies.
No third party has the real item pool, the real difficulty calibration, or the real adaptive algorithm. Every non-ETS test is therefore a good-faith reconstruction. That means two things. First, the question difficulty is somebody's judgement rather than a measured statistic. Second, and more importantly, the raw-to-scaled conversion is an estimate, so the number at the end carries the test-maker's assumptions baked in.
| Source | Real ETS scoring? | What it is genuinely good for | How to read the score |
|---|---|---|---|
| ETS PowerPrep | Yes, the real algorithm and retired official items | The single trustworthy score estimate, and interface rehearsal | Treat a clean timed sitting as your best available forecast |
| Third-party full-length mocks | No, an estimated conversion | Stamina, pacing practice, exposure to more question types | Watch the trend across several, not any single number |
| Our free mock tests | No, an estimated score plus your exact official ETS percentile | Repeat timed practice with a worked solution for every question | Use for volume and diagnosis between the two official tests |
| Untimed question sets and drills | Not applicable, no score | Learning, which is what actually raises your ceiling | Never convert drill accuracy into a predicted score |
The widely repeated claim that a particular company runs harder or easier than the real test is partly true and partly folklore, and it varies by section and by edition. What is reliably true: the further a test is from the official scoring, the more you should read its output as a direction rather than a destination. Improving from 305 to 315 across three third-party mocks is meaningful. The 315 itself is not.
Practical rule: use third-party and free mocks for reps and diagnosis, use official tests for measurement, and never mix the two when you are deciding whether you are ready.
What score should I actually expect on test day?
The forum version of this question is usually posted as a pair of numbers with no context, which is why it never gets a good answer. A mock score only means something once you attach the conditions to it. Here is a framework that produces an honest expectation instead of a hopeful one.
- **Throw out every score that fails the conditions test.** Untimed, interrupted, split across days, or a repeat of questions you had seen. These are not data. Most people find this step removes half their scores.
- **Weight official over third-party.** A clean timed PowerPrep 2 outranks four third-party mocks. If they disagree, believe the official one.
- **Take the trend, not the peak.** People remember their best mock and forget it was the outlier. Use your recent average of valid sittings, not your record.
- **Subtract for missing stamina.** If your valid scores came from tests you took in pieces, expect the real result to land below them.
- **Attach an error bar.** Even a clean official test is an estimate. Plan around a range of a few points per section rather than a single figure.
Applied honestly, this usually produces a narrower and slightly lower expectation than people start with, which is exactly its value. It is far better to discover that in the week before the test, when rescheduling is still an option, than on the score screen.
Once you have a realistic expected range, check it against what you actually need rather than against a number you saw on a forum. Our score calculator converts a scaled score to its official percentile, and what is a good GRE score covers what different fields genuinely expect. Many people discover their range already clears their programmes and that the anxiety was unearned.
How to improve your GRE Verbal score
Verbal is where the practice-to-real gap is usually widest, because so much of it rests on vocabulary, and vocabulary is the easiest thing to accidentally fake in practice. If you took a mock, reviewed the answer key, then took the same mock again, you learned that test rather than the language.
- **Words are the lever, and there is no way around it.** Text Completion and Sentence Equivalence are decided largely by whether you know the word. Reading Comprehension is easier when dense sentences are not also unfamiliar ones. Work through a structured list rather than collecting words at random from mocks, using the vocabulary trainer.
- **Learn words in context, not as pairs.** A word you can only recognise beside its definition will not survive a sentence written to disguise it. This is why our word pages carry a usage sentence rather than a bare gloss.
- **Attack Sentence Equivalence through the sentence, not the choices.** Predict the missing word before reading the options, then find the two choices that match your prediction. Going choice by choice is how near-synonym traps catch people. The Sentence Equivalence guide works this through.
- **Treat Text Completion as a logic puzzle.** The blank is determined by a pivot word somewhere else in the sentence. Find the pivot first. The Text Completion system covers the method.
- **For Reading Comprehension, justify every answer from the passage.** If you cannot point at the line that makes it right, you guessed. See the Reading Comprehension guide.
The fastest honest Verbal gain for most people is vocabulary volume plus a strict habit of justifying answers from the text. Both compound, and neither depends on your authority to interpret a passage cleverly.
How to improve your GRE Quant score
Quant gaps split cleanly into three causes, and the fix for one does nothing for the others. Diagnose before you study. Go back through your last two valid mocks and label every miss.
| If your misses are mostly | The cause is | What actually fixes it |
|---|---|---|
| Questions you understood but got wrong through a slip | Careless execution, not knowledge | Write out every step, stop doing arithmetic in your head, and check what the question asked for before selecting |
| Questions you ran out of time for | Pacing and question triage | Practise abandoning a question at a fixed limit. A skipped hard question costs less than three rushed easy ones |
| Questions you did not know how to start | A genuine content gap | Targeted topic work, plus formula recall. See the GRE math formulas cheat sheet |
| Questions where you fell for a plausible wrong answer | Trap recognition | Study the trap patterns directly in GRE Quant traps and how to beat them |
The most common misdiagnosis is treating careless errors as a content gap. People whose misses are slips go and relearn geometry, which does not help, when the actual fix is writing more down. Drill against the diagnosis with Quant practice by type and difficulty, where every question carries a worked solution so you can see exactly where your reasoning diverged.
One more honest point about the ceiling. Because the GRE is section-adaptive, the top of the Quant scale is only reachable through the harder second-section pool. If you are consistently landing in the easier second section, the work is not polish, it is accuracy on the first section.
When the honest answer is that you need another attempt
Sometimes the gap between practice and reality is not a fixable execution problem, and the useful thing is to say so. If your valid, timed, first-exposure scores were always below your target, then test day did not go wrong. It reported accurately, and the target needs either more study time or a re-think.
That is a decision with real rules attached: how soon you can sit again, what schools see, and what it costs. We cover that fully in how many times you can take the GRE, which includes a retake decision framework and the ScoreSelect rules. This section is only here to make one point: decide from valid data, not from your best mock.
Do not book a retake for next month while keeping exactly the same study routine. A retake without a diagnosis usually reproduces the score. Work out which of the causes above applied to you first.
A practical checklist before your next full-length test
- Have I used PowerPrep 1 for interface rehearsal and kept PowerPrep 2 unopened?
- Am I sitting this full length, in one block, at roughly my real test time?
- Are these questions genuinely new to me?
- Is the timer running for every section, with only the scheduled break?
- Am I writing my work down rather than holding it in my head?
- After scoring, will I label every miss as careless, timing, content, or trap?
- Am I comparing against my recent average of valid sittings, not my best ever score?
If you can answer yes to all seven, the number you get is worth trusting. Take a full-length mock test under exactly those conditions and you will have a real reading of where you stand, and a diagnosis you can act on.