The Sections That Earn Their Place
What to include in a test plan, and what each section is actually for. Sizing decides how many of these test plan sections you write.
Test Scope: What Is In, and the Harder Half, What Is Out
Scope has two halves and almost everybody writes only one.
The in-scope half is easy and rarely argued with: features, platforms, browsers, devices, integrations, and the kinds of testing you will run.
The out-of-scope half is where the value is. This is the list of things that will not be tested on this release, and it is the only part of the plan that transfers risk explicitly. Writing performance testing is out of scope for this release and having product agree is a different situation from nobody mentioning it.
Be specific. Not testing legacy report module is useful. Testing the main features is not.
The kinds of testing worth naming in or out: functional, regression, performance, security testing, accessibility, compatibility, and user testing. Each one is either in or out, and saying nothing means out in practice and in expectation.
One more line worth adding: what proportion of the in-scope work is automated versus manual. Our post on automated versus manual testing covers how to decide, and the plan should record the decision rather than re-argue it.
Risk: What You Are Protecting, Ranked
Risk-based testing means testing the things that matter most, which requires saying what matters most.
List what would actually hurt, and not a generic risk register but the specific things about this release: a payment path that has changed, an integration nobody has touched in a year, a migration that runs once, a feature the largest customer asked for.
Rank them, because the ranking is what you defend when the schedule compresses and somebody asks what can be dropped.
This section does more work than its length suggests. A ranked risk list turns we need two more weeks into without two more weeks, these three things go untested, which is a conversation somebody can actually make a decision about.
Entry and Exit Criteria, and the One Everybody Forgets
Entry criteria say what has to be true before testing starts. The build deploys, the environment is up, unit tests pass, and the feature is actually complete rather than nearly.
Exit criteria say what has to be true before you are done. No open critical defects, agreed coverage of the ranked risks, regression suite passing, and somebody named has signed.
Both are commitments by other people, which is why they belong in a plan rather than in a QA checklist. Entry criteria in particular are a polite way of saying testing does not start when the calendar says so, it starts when the build is testable.
Suspension and Resumption Criteria: When to Stop Testing
The section almost every plan omits, and the one that saves a week.
Suspension criteria say when testing stops before it is finished. The build is too broken to proceed, the environment is unavailable, a blocking defect makes further testing meaningless, or more than some number of tests are failing for the same upstream reason.
Resumption criteria say what has to happen before it restarts.
Without these, a team keeps testing a broken build because stopping feels like giving up, and produces a defect list that is really one defect reported forty times. Agreeing the rule in advance makes stopping a procedure rather than a judgement call somebody has to defend.
Test Environment and Test Data
These two cause more delay than any other part of testing, and they are commitments by other teams.
For the test environment: which one, who provides it, when it is available, how close to production it is, and what is different. The differences matter, because an environment with a tenth of the data will not show you a query that degrades at scale.
For test data: what is needed, who creates it, whether it can be regenerated, and what the rule is for personal data. We will use a copy of production is a decision with legal consequences and it belongs in writing.
Both sections should name a person and a date, because an environment nobody owns arrives late, and the plan is where that gets prevented.
Defect Management: Severity, Priority and Who Decides
Three things to settle before the first defect is raised.
Severity and priority are different. Severity is how badly it breaks, and priority is how soon it gets fixed. A cosmetic issue on the login page can be low severity and high priority. Teams that conflate them argue about single-number labels forever.
Define the levels. Write one line per severity level saying what qualifies. Without that, severity is assigned by whoever is most annoyed.
Name who decides. Somebody has to rule on whether a defect blocks the release, and it cannot be the person who found it or the person who wrote it. Triage with a named owner ends more arguments than any process document.
Roles, Responsibilities and the Name Against Each One
Roles in a plan are worthless without names. The development team will provide test data commits nobody.
That is what roles and responsibilities mean in practice, so list the real activities and put a person against each: who writes the tests, who runs them, who maintains the environment, who provides the data, who triages defects, who approves the exit criteria, who can stop the release.
That last one is the important one, because if nobody can stop a release, the exit criteria are advisory, and the plan is a wish.
Where the names do not exist because the capacity is not there, that is a resourcing finding rather than a documentation problem, and our post on QA and automation testing staff augmentation covers the honest version of that conversation.
Schedule, Test Estimation and the Date You Will Be Asked For
You will be asked how long testing takes, and whatever you write will be treated as a commitment, so write it carefully.
Estimate from the work rather than from the gap in the calendar: how many test cases, how long a full regression cycle takes, how many cycles you expect, and how much time defect verification absorbs. Two full cycles plus fixing time is a more defensible shape than a single number.
State the assumptions the estimate depends on, in the plan, next to the estimate. The build arrives on this date, the environment is available from this date, and defects are fixed within a working day. Every one of those is somebody else's commitment, and when one slips, the estimate moves with it — but only if you wrote it down first.
And give the estimate a shape rather than a point. Eight working days if the build is stable, twelve if we see the defect rate we saw last release is more useful than a single number, and more honest.