Building a grading prompt with the AI Rubric Code Reviewer, running it in the AI assistant you already use, and turning the review into your next round of practice. Screenshots sit behind the same "More detailed info" toggles the other field guides use.
Short version: "is my code good?" gets you vague praise, while a fixed rubric gets you scores you can compare from one week to the next.
Students in a university course get their problem sets graded by a teaching assistant against a rubric. Self-taught learners usually get either nothing or an AI chat that says the code "looks great, with a few small suggestions". This tool closes that gap. It wraps your code in a strict grading prompt, and your own assistant does the grading.
The prompt fixes three things a casual chat leaves loose: what gets judged (the same five dimensions every time), how scores are set (written definitions of each score from 1 to 4, backed by line-number evidence), and what happens next (hints before answers, so you still do the fixing). Because the grading runs in your assistant, your code never touches Deliberate Learners.
Build your grading prompt. Everything below happens inside rubric-code-reviewer.html, entirely in your browser.
One form: what course you're in, how hard to grade, your code, and where you'll paste the result.
Only the code box is required. The tool remembers what you entered in this browser, so you can close the tab and come back after fixing your code. The ↺ button clears the form, and Try an example above the code box loads a buggy binary search if you want to see a review before using your own work.
A course preset tells the grader what that course cares about. Use General for anything else.
Picking a preset fills in the usual language, which you can still change. It also adds that course's priorities to the grading instructions:
| Preset | What the grader weighs extra |
|---|---|
| Harvard CS50 | Every case in the spec, memory safety in C, CS50 style |
| Berkeley CS61A | Abstraction, recursion, higher-order functions, using the technique being taught |
| MIT 6.006 | Correctness argument and running-time analysis as much as working code |
| MIT 6.100A | Clear decomposition and naming at an introductory level |
| Stanford CS106B | Recursive structure, abstract data types, choice of container |
| General | Nothing extra: the five rubric dimensions only |
University is the everyday setting. Use Strict before an exam or interview, and Intro course while you're still learning the basics.
Keep the same level from one submission to the next. Scores only compare meaningfully when the standard stays the same.
The assignment is what the code gets graded against. Without it, the grader has to guess what "correct" means.
Paste the full problem statement, including any stated limits such as "must run in O(log n) time". Those are exactly what the Correctness and Efficiency scores are checked against. If you leave it out, the grader states the assumption it made at the top of the review.
Paste your code as-is. The tool adds line numbers to the prompt so every piece of feedback can point to an exact line. Your code box stays unchanged. For sharper feedback, grade one function or one file at a time. The tool warns you if the code is very long.
Optional, but it makes the review sharper. The grader reconciles your test results with its own hand traces.
If your course has an official autograder, run it first: check50 and
style50 for CS50, ok for CS61A. Paste the output here. Your
own test runs work too. Real test results beat a hand trace, and when the two
disagree, the prompt asks the grader to explain why.
Leave "Include corrected code" off to get hints only. Fixing it yourself is where the learning happens.
With the box off, the prompt tells the grader not to write corrected code, and to give three-step hints instead. Turn it on only when you've already fixed the code yourself and want to compare, or when you're completely stuck. Even then, the prompt asks for the smallest possible change, with each changed line explained.
The assistant you pick only changes which chat the Copy & open button opens. The prompt itself works in any capable assistant. Unlike the Syllabus Reverse-Engineer, this one doesn't need web search.
Build my grading prompt shows the full text. Copy & open copies it and opens a new chat, where you paste it.
Your code makes the prompt too long to send pre-filled, so Copy & open always opens a blank chat and puts the prompt on your clipboard. Just paste. It's worth skimming the prompt once to see the rubric you're being graded against.
Before you paste lists the next steps. With a course that has an official autograder, it reminds you to run it first:
Run it in your assistant. Any capable assistant works, and no web search is needed.
An earlier conversation, including an earlier review of the same code, can color the grade. A fresh chat keeps each review independent.
Don't trim the rubric. The written definitions of each score are what keep grades consistent from one review to the next.
The traces show the code running on real inputs. That's where bugs become visible, and where you can check the grader's work.
Each problem comes with a first hint. Try it, and reply next hint only when you're genuinely stuck.
The prompt tells the grader not to raise a score just because you push back. Point to the line it misread, and it will reconsider.
Read the review. The same sections come back every time, in this order.
| Section | What it tells you |
|---|---|
| Scorecard | A score from 1 to 4 on each of the five dimensions, each with a one-line reason citing line numbers, plus a total out of 20. |
| Traces | 3 to 5 inputs, including edge cases, run by hand: expected result, actual result, pass or fail. |
| Isolated Skill Mastery | What your code proves you can already do, with the lines that show it. |
| Structural Blind Spots | Habits missing from how you approach problems like this, most costly first. |
| Root Cause Diagnostics | Each problem's symptom, its cause from a fixed list, and why. |
| Hints | A first hint for each of the three biggest problems. More on request. |
| Corrected version | Only if you ticked "Include corrected code". |
| Next rep | One exercise, under 30 minutes, aimed at your weakest dimension. |
Every score is set against the same written definitions, so a 3 this week means the same as a 3 last week.
| Dimension | 4 means | 1 means |
|---|---|---|
| Correctness | Right on every traced input | Doesn't work, or doesn't address the assignment |
| Edge cases | Every boundary and invalid input the assignment implies is handled | Only the happy path works |
| Efficiency | Optimal or clearly appropriate complexity | Impractically slow for realistic inputs |
| Structure | Clear names, sensible decomposition, idiomatic | Very hard to read or maintain |
| Testing | Evidence of testing that includes edge cases | No sign the code was ever run |
Testing scores low if you paste code alone. Pasting your test output, or including your tests or asserts in the code box, is how you earn it.
The cause matters more than the bug. Each one calls for a different kind of practice.
| Root cause | What to practice |
|---|---|
| Concept gap | Go back to the lecture or reading on that idea before writing more code. |
| Careless slip | You knew better. Slow down, and test before submitting. |
| Missed edge case | Write down empty, single-element, and boundary inputs before you start coding. |
| Wrong algorithm choice | Compare approaches and their running times before you commit to one. |
| Misread the spec | Restate the assignment in your own words before starting. |
| Language or tooling quirk | Look up the behavior in the language docs and write a two-line test to confirm it. |
Keep a running tally across reviews. If the same cause keeps coming back, that habit is worth more of your practice time than any single bug.
Close the loop. One review is feedback. Repeating the cycle is practice.
AI graders can make mistakes when tracing code by hand. When a trace and your own run of the code disagree, trust the run. The official autograder always wins over this review.
Match the symptom to the likely cause before starting over.
| Symptom | Likely cause and fix |
|---|---|
| It gave the corrected code anyway | Some assistants override instructions. Reply "hints only, no code", or start a new chat and paste again. |
| Everything scored 4 | The standard is too loose, or the assignment was missing. Add the assignment, or switch to Strict. |
| A trace says the code fails, but it works when you run it | A hand-tracing mistake. Reply with the actual output, and ask it to re-trace that input. |
| It assumed the wrong language | Fill in the Language field and rebuild the prompt. |
| The review stops partway | It hit its length limit. Reply continue. |
| Feedback is vague for long code | Too much at once. Grade one function or file per review. |