Reference

How scoring works

Every grade comes from a published, versioned rubric. You can redo any score by hand from the findings in the report.

The sum

A server starts at 100. Each finding subtracts a penalty by severity:

SeverityPenalty
critical40
high20
medium8
low3
info0
score = clamp(100 - sum of penalties, 0, 100)

Each rule counts at most once per tool, so one systemic problem cannot subtract twenty times.

One critical finding means F

Any critical finding caps the score at 39, which is an F, whatever else the server gets right. A server that reads your SSH key on startup does not deserve a D for good manners elsewhere. The grade line says capped by a critical finding when this happens.

Grades

GradeScore
A90 to 100
B75 to 89
C60 to 74
D40 to 59
F0 to 39

A worked example

A remote server that lists its tools without authentication (high, minus 20) and carries a base64 blob in one description (medium, minus 8) scores 100 minus 20 minus 8, which is 72, a C. Add a decoy credential read (critical) and the cap holds it at 39, an F.

Versions

Every report records the MCPsight version and the rubric version that produced it. When the rubric changes, its version goes up and old grades keep their meaning. Nothing is silently re-graded. The full rubric is in docs/rubric.md.