← all posts

Turning Coding Test Results into Curriculum-Level Learning Diagnostics

‘Coding tests leave scores, but make it difficult to identify what students don't understand.
This service uses NCIC informatics achievement standards to guide AI-generated questions and turn grading results into learning diagnostics.
It aims to provide a diagnostic framework for roughly 70,000 coaching centers in India that lack a standardized curriculum.’

That's the idea in brief. Its most important phrase isn't ‘generate questions with AI.’ It's ‘turn grading results into learning diagnostics.’


A score of 80 is accurate, but not sufficient

Suppose two students both score 80 on the same coding test.

One doesn't understand how conditions behave inside loops. The other struggles to trace and fix program errors. Both reports say 80. But the explanations and practice they need in the next lesson are entirely different.

A score tells us how much a student got right. It doesn't necessarily tell us what they don't understand or what to teach next.

Coding tests expose this clearly. An autograder can accurately count passed cases out of ten. It takes separate interpretation to determine whether failures came from array boundaries, nested-condition execution order, or poor function decomposition.

Grading evaluates code's result. Teaching needs to understand why that result occurred.

That gap was the problem I focused on for the 8th Education Public Data AI Utilization Competition.

I didn't start with ‘let's generate coding questions with AI.’ Asking generative AI for ten Python-loop questions is already easy. What's missing is a definition of what each question assesses, what a wrong answer suggests, and how that should change the next lesson.

More questions don't automatically produce better diagnostics.

Question generation without criteria only automates the labor of writing questions.


I wanted to change the unit of data a test leaves behind

Ordinary test data looks like this:

Student A
Score: 80
Incorrect: questions 3 and 7

That's sufficient for grading, but not for planning a lesson. It doesn't connect questions 3 and 7 to the concepts they measured.

I wanted a result like this:

Student A
Score: 80

[9정03-06] Logical operations and nested control structures
- 2 correct answers out of 4 related questions
- Repeated errors tracing execution order in nested conditions

[9정03-07] Error correction using functions and a debugger
- 3 correct answers out of 3 related questions
- Consistently identifies and corrects errors

Both can come from the same test. Only the second answers what should happen in the next lesson.

The key is not merely attaching an answer to a question. It's recording which learning objective that question provides evidence about.

The required structure is simple:

Curriculum
→ Achievement standard
→ Assessment element
→ Question
→ Student response
→ Evidence by achievement standard
→ Remediation and the next lesson

Accumulating results this way turns question 3 was wrong into educational information: difficulty interpreting nested control structures.


Achievement standards became a shared language for interpreting errors

This is where data from NCIC, Korea's National Curriculum Information Center, became necessary.

NCIC provides curriculum documents from Korea and other countries and supports navigation through structured curriculum information. The student-assessment portal lets educators find achievement and assessment standards by school level and curriculum, alongside assessment tools and resources for using results.

An achievement standard isn't merely a chapter title.

Rather than naming topics like loops, functions, or lists, it describes what students should understand and be able to do after learning.

For example, Korea's revised 2022 middle-school informatics curriculum separates solving problems with sequential data structures, using logical operations and nested control structures, and analyzing and correcting errors with functions and debuggers into distinct standards.

That distinction matters when interpreting results.

Memorizing for syntax isn't the same as solving a problem with iteration. Choosing correct code from alternatives isn't the same as using a debugger to trace an error's cause.

Standards shift assessment from checking which syntax students learned toward observing what they can actually do.

The Ministry of Education's process-oriented assessment likewise gathers evidence of student growth against curriculum standards and uses it for feedback. My proposal tried to make that principle easier to apply in coding academies and other nonformal education settings.

Achievement standards weren't just a prompt for AI. They were the data criteria connecting questions, responses, and feedback.


The same score should yield different diagnoses

The following is a hypothetical example, not actual learner data.

Student Score Main errors Interpretation by standard Next learning activity
A 80 Tracing nested-condition execution Insufficient evidence of understanding logical operations and nested control structures Manually trace execution order
B 80 Analyzing and correcting a faulty function Insufficient evidence of function decomposition and debugging Practice debugging with breakpoints and variable values

In a conventional report, the students look identical.

Student A: 80
Student B: 80

At the achievement-standard level, they differ.

Student A: needs practice interpreting control structures
Student B: needs practice tracing and correcting errors

That difference is the diagnosis.

Diagnosis isn't generating a long, convincing-sounding explanation. It's distinguishing learning states hidden beneath the same score and changing the next action accordingly.


Public data as product structure, not reference material

Public-data proposals sometimes use data only to justify slides—market size or problem severity—while the product would work identically without it.

Remove NCIC data from this proposal, and the central product structure disappears too.

Each question needs these fields:

Question ID
Achievement-standard ID
Target performance
Difficulty
Correct answer or rubric
Student response
Learning evidence provided by that response

One wrong answer doesn't justify declaring that a student ‘doesn't understand loops.’ We need to check the intended standard, responses to other questions linked to it, and whether the question itself contains errors or ambiguity.

Connecting questions and responses through standard IDs enables aggregation:

  • Standards an individual repeatedly struggles with.
  • Standards where a whole class concentrates errors.
  • Questions whose observed success rate differs from intended difficulty.
  • Areas that don't improve after teaching.
  • Students with the same score but different remediation needs.

NCIC wasn't a citation at the back of a report. It was a schema that made test data interpretable.


AI's role was reducing translation effort, not making final judgments

Even if standards-based assessment is educationally sound, repeatedly finding standards, decomposing assessment elements, and writing questions and rubrics is burdensome.

That's where AI was useful.

AI can assist with:

  1. Decomposing standards into age-appropriate assessment elements.
  2. Drafting questions and rubrics for those elements.
  3. Explaining why a question aligns with a standard.
  4. Organizing responses into evidence by standard.
  5. Drafting remediation activities and feedback for teachers.

AI-generated questions shouldn't enter a test without review, and AI shouldn't make the final judgment about student attainment.

Research on LLM-generated educational questions emphasizes quality control for curriculum alignment and errors. UNESCO also recommends keeping human oversight and final responsibility with educators rather than replacing teachers' responsibilities with AI.

The workflow I designed therefore looks like this:

Eight-step workflow connecting curriculum standards, assessment, evidence, and remediation

AI doesn't replace the standards. It reduces repetitive work between standards and actual assessment.

Standards set the direction. AI lowers the cost of moving in that direction.

The distinction mattered. This should help teachers write questions faster and read results more specifically, not classify students as simply ‘good’ or ‘bad.’


Why global coding education?

I considered India's private coaching market as the first candidate for application.

India's 2025 Comprehensive Modular Survey on Education, published by the Ministry of Statistics and Programme Implementation, reported that 27.0% of enrolled students had used or were using private coaching or tuition during the academic year: 30.7% in urban areas and 25.5% in rural areas.

Private coaching is therefore not merely a marginal service. It's part of many students' actual learning paths.

But exporting Korean standards unchanged wouldn't be appropriate. Curricula, assessment philosophies, and learning contexts differ across countries.

I used NCIC not because Korea's curriculum is the universal answer, but because it offered a structured starting point for showing what standards-linked assessment data makes possible.

A global service would extend the structure as follows:

Korea: NCIC achievement standards
India: national/state curricula or institutional standards
Other countries: local official curricula
Private institutions: their own curricula in the same structured format

The criteria's content should change. The connection standard → question → response → diagnosis should remain.


The recognition wasn't only for AI question generation

The idea received a Bronze Award at the 8th Education Public Data AI Utilization Competition.

An award doesn't validate educational effectiveness. A competition judges the problem, proposal, and potential use of public data and AI. It isn't an educational experiment demonstrating learning gains.

I think three things made the proposal convincing.

First, the problem was clear.

Coding tests leave scores but insufficient evidence of what students don't understand.

Second, public data wasn't decoration.

NCIC standards formed the data structure connecting question generation with result analysis.

Third, AI's purpose wasn't another item on a feature list.

AI reduced the work of translating standards into questions and feedback rather than merely adding a chatbot.

The differentiator wasn't how impressive AI-generated sentences sounded.

It was trying to make the data left after a test more useful for teaching.


What's needed now is validation data, not a better presentation

To turn this into a product, the first question isn't whether AI can generate questions. We can already demonstrate that.

The questions to validate are harder:

Validation question Evidence to measure
Do generated questions assess their intended standards? Expert alignment judgments and inter-rater agreement
Does teacher workload actually fall? Time from drafting to approval, revision and rejection rates
Are diagnostics specific enough for the next lesson? Teacher usefulness ratings and time to select remediation
Can equal scores yield distinct profiles? Profile differentiation among students with the same score
Does targeted remediation help? Reassessment against the same standard before and after remediation
Does alignment survive language and country changes? Alignment after translation and local teacher revision rates

A minimum validation could begin like this.

Choose five standards and generate ten questions per standard, for 50 total. At least two educators independently judge alignment, answer clarity, and suitability for the student level. Record proportions usable unchanged, usable after revision, and requiring rejection.

Then collect student responses and check whether equal total scores actually hide different standard profiles. Finally, measure whether teachers can choose the next remedial activity faster and more specifically using the diagnostics.

Only with that evidence can we say:

  • AI generates aligned questions with a measured degree of reliability.
  • Teacher review time actually fell.
  • Diagnostics enabled more specific remediation than total scores alone.
  • Targeted remediation improved reassessment against a particular standard.

For now, these are hypotheses to validate, not results.

The award recognized potential. Actual usage must demonstrate the product's value.


A test should start the next lesson, not end with a number

The biggest change in my thinking wasn't about AI.

At first, question-generation speed seemed important. Looking deeper, what mattered more was what remained after the test.

Leave only a score, and the test summarizes the student and stops.

Leave evidence by standard, and the test begins the next lesson.

80

Turn that one line into:

You understood control structures,
but repeatedly struggled during debugging.
Next, practice tracing errors step by step.

That is educational information.

This transformation is the core of the standards-based AI assessment system I proposed.

AI doesn't need to replace teaching. It should reduce repetitive work so teachers can understand students more specifically. Public data shouldn't end as numbers making a presentation persuasive; it should give educational meaning to student responses.

What we wanted to change wasn't question-generation speed.

Good assessment doesn't stop at ranking students. It leaves information that changes the next lesson.