Review time and comment patterns reveal data quality issues.
Our data validation process is based on total review time, time spent on each review, the number of comments left on each project, and more. Most variables followed a Gaussian distribution, allowing us to use various statistical tools similar to those that identified cheating sumo wrestlers or overlooked baseball players. Based on predicted values from these metrics, we can identify problematic reviewing patterns. This approach mirrors data quality assurance in machine learning, where high-quality labeled data leads to better model performance.