Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

RAG Model Tested for Skin Lesion Grading

RAG Model Tested for Skin Lesion Grading

Semantic Scholar·Sunday, August 9, 2026
  • •Claude 4.5 Opus RAG system graded 60 cSCC and 67 melanocytic nevus cases
  • •cSCC grading reached 90.0% concordance and κ = 0.85 against dermatopathologist references
  • •Nevus grading reached 44.8% concordance, with no severely dysplastic case concordantly classified
  • •Claude 4.5 Opus RAG system graded 60 cSCC and 67 melanocytic nevus cases
  • •cSCC grading reached 90.0% concordance and κ = 0.85 against dermatopathologist references
  • •Nevus grading reached 44.8% concordance, with no severely dysplastic case concordantly classified
  • •Claude 4.5 Opus RAG system graded 60 cSCC and 67 melanocytic nevus cases
  • •cSCC grading reached 90.0% concordance and κ = 0.85 against dermatopathologist references
  • •Nevus grading reached 44.8% concordance, with no severely dysplastic case concordantly classified
  • •Claude 4.5 Opus RAG system graded 60 cSCC and 67 melanocytic nevus cases
  • •cSCC grading reached 90.0% concordance and κ = 0.85 against dermatopathologist references
  • •Nevus grading reached 44.8% concordance, with no severely dysplastic case concordantly classified

Joshua Mijares, Eric Gan, Neil K. Jairath and colleagues reported in Cancers in 2026 that a retrieval-augmented generation system (retrieving relevant documents before answering) using Claude 4.5 Opus was tested for histopathologic grading of skin lesions. The study evaluated whether RAG-assisted AI could support grading when dermatopathologists face inter-observer variability in cutaneous squamous cell carcinoma differentiation and melanocytic nevus dysplasia.

The researchers built two isolated RAG pathways, each grounded with ChromaDB vector retrieval (searching text by mathematical similarity) from task-specific literature. One pathway graded 60 cSCC cases by differentiation, while the other graded 67 melanocytic nevus cases by dysplasia. Each case was analyzed three times, and the majority result became the final grade.

For cSCC, the LLM reached 90.0% concordance with reference dermatopathologist grading, with a 95% CI of 79.9–95.3%. Inter-rater reliability was excellent, with Cohen’s kappa (agreement beyond chance) of κ = 0.85 and a 95% CI of 0.74–0.96.

For nevi, concordance was 44.8%, with a 95% CI of 33.5–56.6%, and agreement was not statistically distinguishable from chance, with κ = 0.072 and a 95% CI of −0.10–0.24. Discordant nevus classifications skewed toward moderately dysplastic results, accounting for 65.7% of LLM classifications versus 41.8% of reference classifications. No severely dysplastic case was concordantly classified, at 0/6.

The authors concluded that the identical RAG architecture performed differently by task: high for cSCC differentiation and low for nevus dysplasia grading. They said the pattern parallels reported human inter-rater reliability and supports task-specific validation before clinical or educational use.

Joshua Mijares, Eric Gan, Neil K. Jairath and colleagues reported in Cancers in 2026 that a retrieval-augmented generation system (retrieving relevant documents before answering) using Claude 4.5 Opus was tested for histopathologic grading of skin lesions. The study evaluated whether RAG-assisted AI could support grading when dermatopathologists face inter-observer variability in cutaneous squamous cell carcinoma differentiation and melanocytic nevus dysplasia.

The researchers built two isolated RAG pathways, each grounded with ChromaDB vector retrieval (searching text by mathematical similarity) from task-specific literature. One pathway graded 60 cSCC cases by differentiation, while the other graded 67 melanocytic nevus cases by dysplasia. Each case was analyzed three times, and the majority result became the final grade.

For cSCC, the LLM reached 90.0% concordance with reference dermatopathologist grading, with a 95% CI of 79.9–95.3%. Inter-rater reliability was excellent, with Cohen’s kappa (agreement beyond chance) of κ = 0.85 and a 95% CI of 0.74–0.96.

For nevi, concordance was 44.8%, with a 95% CI of 33.5–56.6%, and agreement was not statistically distinguishable from chance, with κ = 0.072 and a 95% CI of −0.10–0.24. Discordant nevus classifications skewed toward moderately dysplastic results, accounting for 65.7% of LLM classifications versus 41.8% of reference classifications. No severely dysplastic case was concordantly classified, at 0/6.

The authors concluded that the identical RAG architecture performed differently by task: high for cSCC differentiation and low for nevus dysplasia grading. They said the pattern parallels reported human inter-rater reliability and supports task-specific validation before clinical or educational use.

Read original (English)·Aug 6, 2026
Healthcare#retrieval augmented generation#claude 4 5 opus#chromadb#histopathology#cutaneous squamous cell carcinoma#melanocytic nevi#cohens kappa#dermatopathology