Reference
How scoring works
Every grade comes from a published, versioned rubric. You can redo any score by hand from the findings in the report.
The sum
A server starts at 100. Each finding subtracts a penalty by severity:
| Severity | Penalty |
|---|---|
| critical | 40 |
| high | 20 |
| medium | 8 |
| low | 3 |
| info | 0 |
score = clamp(100 - sum of penalties, 0, 100)
Each rule counts at most once per tool, so one systemic problem cannot subtract twenty times.
One critical finding means F
Any critical finding caps the score at 39, which is an F, whatever else the server gets right. A server that reads your SSH key on startup does not deserve a D for good manners elsewhere. The grade line says capped by a critical finding when this happens.
Grades
| Grade | Score |
|---|---|
| A | 90 to 100 |
| B | 75 to 89 |
| C | 60 to 74 |
| D | 40 to 59 |
| F | 0 to 39 |
A worked example
A remote server that lists its tools without authentication (high, minus 20) and carries a base64 blob in one description (medium, minus 8) scores 100 minus 20 minus 8, which is 72, a C. Add a decoy credential read (critical) and the cap holds it at 39, an F.
Versions
Every report records the MCPsight version and the rubric version that produced it. When the rubric changes, its version goes up and old grades keep their meaning. Nothing is silently re-graded. The full rubric is in docs/rubric.md.