Research1w ago
Reliability, validity, and diagnostic evidence for multi-model LLM short-answer scoring
A study evaluates multi-model LLM scoring using GPT, DeepSeek, and Qianwen for short-answer educational assessment, showing high reliability, strong validity…
#DeepSeek#LLM#educational assessment#reliability#short-answer scoring#GPT