Field guide

From pasted code to a graded scorecard

Building a grading prompt with the AI Rubric Code Reviewer, running it in the AI assistant you already use, and turning the review into your next round of practice. Screenshots sit behind the same "More detailed info" toggles the other field guides use.


Optional · About this tool

Why grade against a fixed rubric?

Short version: "is my code good?" gets you vague praise, while a fixed rubric gets you scores you can compare from one week to the next.

More detailed infoLess info

Students in a university course get their problem sets graded by a teaching assistant against a rubric. Self-taught learners usually get either nothing or an AI chat that says the code "looks great, with a few small suggestions". This tool closes that gap. It wraps your code in a strict grading prompt, and your own assistant does the grading.

The prompt fixes three things a casual chat leaves loose: what gets judged (the same five dimensions every time), how scores are set (written definitions of each score from 1 to 4, backed by line-number evidence), and what happens next (hints before answers, so you still do the fixing). Because the grading runs in your assistant, your code never touches Deliberate Learners.

Part 1

Build your grading prompt. Everything below happens inside rubric-code-reviewer.html, entirely in your browser.

01

Open the tool

One form: what course you're in, how hard to grade, your code, and where you'll paste the result.

More detailed infoLess info

Only the code box is required. The tool remembers what you entered in this browser, so you can close the tab and come back after fixing your code. The ↺ button clears the form, and Try an example above the code box loads a buggy binary search if you want to see a review before using your own work.

The AI Rubric Code Reviewer form, empty
The whole form on first load
02

Pick your course and language

A course preset tells the grader what that course cares about. Use General for anything else.

More detailed infoLess info

Picking a preset fills in the usual language, which you can still change. It also adds that course's priorities to the grading instructions:

PresetWhat the grader weighs extra
Harvard CS50Every case in the spec, memory safety in C, CS50 style
Berkeley CS61AAbstraction, recursion, higher-order functions, using the technique being taught
MIT 6.006Correctness argument and running-time analysis as much as working code
MIT 6.100AClear decomposition and naming at an introductory level
Stanford CS106BRecursive structure, abstract data types, choice of container
GeneralNothing extra: the five rubric dimensions only
Harvard CS50 selected with language C, and Strict grading selected
CS50 preset, language filled in, Strict grading
03

Choose how hard to grade

University is the everyday setting. Use Strict before an exam or interview, and Intro course while you're still learning the basics.

More detailed infoLess info
  • Intro course: expects working, sensibly organized code and doesn't mark you down for skipping advanced techniques.
  • University: the standard a teaching assistant at a top university would apply to a problem set.
  • Strict: exam and technical-interview standard. Any unhandled edge case or avoidable inefficiency costs points, and a 4 should be rare.

Keep the same level from one submission to the next. Scores only compare meaningfully when the standard stays the same.

04

Paste the assignment and your code

The assignment is what the code gets graded against. Without it, the grader has to guess what "correct" means.

More detailed infoLess info

Paste the full problem statement, including any stated limits such as "must run in O(log n) time". Those are exactly what the Correctness and Efficiency scores are checked against. If you leave it out, the grader states the assumption it made at the top of the review.

Paste your code as-is. The tool adds line numbers to the prompt so every piece of feedback can point to an exact line. Your code box stays unchanged. For sharper feedback, grade one function or one file at a time. The tool warns you if the code is very long.

Assignment and code boxes filled with the binary search example
The built-in example: a binary search with two bugs
05

Add autograder or test output

Optional, but it makes the review sharper. The grader reconciles your test results with its own hand traces.

More detailed infoLess info

If your course has an official autograder, run it first: check50 and style50 for CS50, ok for CS61A. Paste the output here. Your own test runs work too. Real test results beat a hand trace, and when the two disagree, the prompt asks the grader to explain why.

Test output box showing one failing pytest test that timed out
A failing test the grader has to account for
06

Choose your assistant, and whether to see corrected code

Leave "Include corrected code" off to get hints only. Fixing it yourself is where the learning happens.

More detailed infoLess info

With the box off, the prompt tells the grader not to write corrected code, and to give three-step hints instead. Turn it on only when you've already fixed the code yourself and want to compare, or when you're completely stuck. Even then, the prompt asks for the smallest possible change, with each changed line explained.

The assistant you pick only changes which chat the Copy & open button opens. The prompt itself works in any capable assistant. Unlike the Syllabus Reverse-Engineer, this one doesn't need web search.

Assistant set to Claude, Include corrected code unticked
Hints only: the recommended default
07

Build and copy the prompt

Build my grading prompt shows the full text. Copy & open copies it and opens a new chat, where you paste it.

More detailed infoLess info

Your code makes the prompt too long to send pre-filled, so Copy & open always opens a blank chat and puts the prompt on your clipboard. Just paste. It's worth skimming the prompt once to see the rubric you're being graded against.

The generated grading prompt with Copy prompt and Copy and open Claude buttons
The finished prompt, ready to copy

Before you paste lists the next steps. With a course that has an official autograder, it reminds you to run it first:

Before you paste instructions, starting with Run ok first
With the CS61A preset, step 1 is running ok

Part 2

Run it in your assistant. Any capable assistant works, and no web search is needed.

01

Start a new chat every time

An earlier conversation, including an earlier review of the same code, can color the grade. A fresh chat keeps each review independent.

02

Paste the whole prompt and send it

Don't trim the rubric. The written definitions of each score are what keep grades consistent from one review to the next.

03

Read the traces before the scores

The traces show the code running on real inputs. That's where bugs become visible, and where you can check the grader's work.

04

Work the hints one step at a time

Each problem comes with a first hint. Try it, and reply next hint only when you're genuinely stuck.

05

Dispute a score only with evidence

The prompt tells the grader not to raise a score just because you push back. Point to the line it misread, and it will reconsider.


Part 3

Read the review. The same sections come back every time, in this order.

SectionWhat it tells you
ScorecardA score from 1 to 4 on each of the five dimensions, each with a one-line reason citing line numbers, plus a total out of 20.
Traces3 to 5 inputs, including edge cases, run by hand: expected result, actual result, pass or fail.
Isolated Skill MasteryWhat your code proves you can already do, with the lines that show it.
Structural Blind SpotsHabits missing from how you approach problems like this, most costly first.
Root Cause DiagnosticsEach problem's symptom, its cause from a fixed list, and why.
HintsA first hint for each of the three biggest problems. More on request.
Corrected versionOnly if you ticked "Include corrected code".
Next repOne exercise, under 30 minutes, aimed at your weakest dimension.
Optional · The rubric

What a 4 and a 1 mean on each dimension

Every score is set against the same written definitions, so a 3 this week means the same as a 3 last week.

More detailed infoLess info
Dimension4 means1 means
CorrectnessRight on every traced inputDoesn't work, or doesn't address the assignment
Edge casesEvery boundary and invalid input the assignment implies is handledOnly the happy path works
EfficiencyOptimal or clearly appropriate complexityImpractically slow for realistic inputs
StructureClear names, sensible decomposition, idiomaticVery hard to read or maintain
TestingEvidence of testing that includes edge casesNo sign the code was ever run

Testing scores low if you paste code alone. Pasting your test output, or including your tests or asserts in the code box, is how you earn it.

Optional · Root causes

What to do about each kind of mistake

The cause matters more than the bug. Each one calls for a different kind of practice.

More detailed infoLess info
Root causeWhat to practice
Concept gapGo back to the lecture or reading on that idea before writing more code.
Careless slipYou knew better. Slow down, and test before submitting.
Missed edge caseWrite down empty, single-element, and boundary inputs before you start coding.
Wrong algorithm choiceCompare approaches and their running times before you commit to one.
Misread the specRestate the assignment in your own words before starting.
Language or tooling quirkLook up the behavior in the language docs and write a two-line test to confirm it.

Keep a running tally across reviews. If the same cause keeps coming back, that habit is worth more of your practice time than any single bug.


Part 4

Close the loop. One review is feedback. Repeating the cycle is practice.

Amber · proceed with care

An AI grade is a second opinion, not the answer key

AI graders can make mistakes when tracing code by hand. When a trace and your own run of the code disagree, trust the run. The official autograder always wins over this review.

Rust · troubleshooting

If the review isn't what you expected

Match the symptom to the likely cause before starting over.

More detailed infoLess info
SymptomLikely cause and fix
It gave the corrected code anywaySome assistants override instructions. Reply "hints only, no code", or start a new chat and paste again.
Everything scored 4The standard is too loose, or the assignment was missing. Add the assignment, or switch to Strict.
A trace says the code fails, but it works when you run itA hand-tracing mistake. Reply with the actual output, and ask it to re-trace that input.
It assumed the wrong languageFill in the Language field and rebuild the prompt.
The review stops partwayIt hit its length limit. Reply continue.
Feedback is vague for long codeToo much at once. Grade one function or file per review.