Writing, Testing and Evaluating the Program
The last three stages of the seven-step process: write the program, test it, evaluate the solution.
Writing the program
Translating the algorithm into Python. If the algorithm is sound, this should be mechanical — and if it isn't mechanical, that's a signal the algorithm wasn't finished.
Practical points:
- Translate one step at a time; resist improving the logic while typing
- Use names that say what they hold (
total, nott) - Comment why, not what — the code already says what
- Run it early and often rather than writing everything then debugging
Testing the program
Testing shows the presence of bugs, never their absence. You cannot test every input; you choose inputs likely to expose failure.
Categories of test case:
| Category | Purpose | Example (average of n numbers) |
|---|---|---|
| Normal | Typical expected input | 5 numbers: 10, 20, 30, 40, 50 |
| Boundary | Edges of valid range | n = 1 |
| Edge | Unusual but legal | n = 0, all values identical |
| Invalid | Should be rejected cleanly | "abc" entered for a number |
| Large | Performance and overflow | n = 100,000 |
Every test needs an expected result decided in advance. Running a program and glancing at the output is not testing — it's watching. If you didn't know what the answer should be before you ran it, you cannot tell whether it's right.
Three kinds of error
1. Syntax errors — the code isn't valid Python. Python refuses to run it and tells you the line. Easiest to fix.
print("Hello" # SyntaxError: missing )
2. Runtime errors — valid code that fails during execution.
x = 10 / 0 # ZeroDivisionError
3. Logic errors — runs fine, gives the wrong answer. The dangerous kind, because nothing complains.
average = a + b + c / 3 # runs, wrong: precedence
That last one connects directly to operator precedence. It produces a number, the program exits cleanly, and the number is wrong.
Evaluating the solution
Testing asks "does it work?" Evaluation asks "is this a good solution?" — a different and broader question.
- Does it actually solve the original problem, or a problem I drifted into?
- Is it efficient enough for realistic input sizes?
- Can someone else read and modify it?
- Does it handle the edge cases we identified at the understanding stage?
- Could a different approach be substantially simpler?
Evaluation is also where you look back at the model's assumptions. If you assumed absent students count as failures, is that still defensible now that you see the results?
Pólya calls this "looking back", and argues it's the stage that actually teaches you something — the others merely get the job done.