OMK — Observe. Measure. Know. Make every knowledge change in your AI application evidence-backed.
-
Updated
Aug 10, 2026 - TypeScript
OMK — Observe. Measure. Know. Make every knowledge change in your AI application evidence-backed.
Measure whether your model judge agrees with human raters. Chance-corrected agreement statistics, bootstrap confidence intervals, and calibration gates for Swift Testing and CI. Zero dependencies.
Measure how much your LLM judges actually agree. Inter-judge agreement metrics for LLM-as-a-judge evaluations.
A reliability and DIF report card for LLM-judge and human-rater scoring instruments.
Tool-agnostic inter-coder reliability (Krippendorff alpha, Cohen/Fleiss kappa) and disagreement adjudication for qualitative coding
Browser-based inter-rater reliability calculator for systematic literature reviews. Computes Krippendorff's Alpha using a pooled coincidence matrix across all (Paper, RQ, Field) units. No installation required — single HTML file, fully client-side. Built for the GenAI Evidence Hub at Learning Data Insights.
Evaluation toolkit for multi-annotator human annotation research with agreement metrics, disagreement analysis, reports, and plots.
Classificazione dei commenti YouTube di Breaking Italy e La Repubblica tramite il modello User Needs (Shishkin & SmartOcto, 2021). Tesi magistrale in Comunicazione, ICT e Media — UniTO.
Reliability, rogue-rater, drift and leakage diagnostics for labelled data — on the raw GoEmotions ratings, 27 of 28 emotions fall below the 0.667 agreement floor.
Add a description, image, and links to the krippendorff-alpha topic page so that developers can more easily learn about it.
To associate your repository with the krippendorff-alpha topic, visit your repo's landing page and select "manage topics."