CodeCognify: can small models diagnose programming skill?
I compare two cognitive diagnostic models, DINA and G-DINA, against three small and two large language models on 93 students, 49 problems, and a Q-matrix mapping problems to 12 skills. The aim is to test whether smaller, more accessible models can support fine-grained skill diagnosis.