An automated, automatically improving framework for describing financial data quality grades at arbitrary levels is proposed, which first train a financial classifier to categorize data into multiple quality grades, with the theoretical capability to support arbitrary grading levels.
Real data often contains errors, which is why data engineers spend a lot of time creating data cleaning pipelines to ensure the best possible data quality. However, it is often difficult to compare the results of different pipelines and decide which pipeline leads to the best results. There are many different metrics that are designed for different use cases, but they often take only a portion of the data into account. There is a lack of universally applicable metrics for measuring data quality that can be used in many different scenarios. That is why in this paper we are presenting TOMME - an initial approach to a universally applicable weighted error-based metric for data quality. This allows the data quality of a dataset to be assessed based on a single score. While a detailed data quality evaluation remains important, the use of a single score enables rapid assessment and automated processing, for example, for optimization algorithms. By using different weights, the score can also be precisely adjusted to the specific use case. That is why we named it TOMME, which stands for"The One Metric Measuring Errors". As the name suggests, it measures errors in the data. It can thus be considered a generalized, weighted form of accuracy.
DeepQual-Web is introduced, a single multimodal deep learning framework for comprehensive quality assessment and optimization of web applications that combines Gradient Boosting Regression and Bidirectional Long Short-Term Memory networks to interpret the performance and reliability attributes of execution logs and system metrics.
I. Alharbi· International Conference on...· 0 citations
This research proposes an Explainable Machine Learning (XML)–based framework to assess software quality by integrating code metrics, defect datasets, and advanced interpretability methods such as SHAP, LIME, and permutation importance.
Nandhini Ravi· International Journal of Mac...· 0 citations
Intelligent document processing (IDP) with vision-language models (VLMs) hinges on confidence scores trustworthy enough to route extractions between automation and human review. Existing document benchmarks are dominated by clean, high-quality samples, leaving low accuracy regions too sparse for calibration assessment. We introduce ConfBench, the first calibration-specific benchmark for key information extraction (KIE), built by applying 20 controlled degradation pipelines to a diverse document set, yielding 1,346 variants and 70K+ entity-level evaluations spanning the full accuracy spectrum. We evaluate four proprietary and three open-weight VLMs under verbalized and log-probability confidence estimation methods across three input modalities, and find: (i) OCR+Image modality results in more accurate confidence estimates; (ii) model capability is the dominant factor: within the Claude family confidence quality scales monotonically with capability, while across families parameter count is a poor predictor; (iii) calibration quality varies widely across models, from near-perfect to severely overconfident, and per-model post-hoc correction rescales these absolute confidence values for threshold-based routing without altering ranking-based operational metrics; and (iv) log-probability with first-token aggregation consistently outperforms mean-token and margin aggregations. We also introduce ECARB, a review-budget metric translating discriminative gains into operational savings. We release ConfBench publicly to enable systematic study of confidence estimators and calibration methods for trustworthy IDP application deployment.
P. Roy, Sujitha Martin, Mohammad Rostami et al.· 0 citations
Grading open-ended technical responses has been a longstanding issue in higher education. Despite advances in automated assessment, existing approaches often rely on holistic scoring, weakly validated annotations, and cognitive alignment, limiting pedagogical reliability and classroom adoption. To facilitate reliable and rubric-based automated evaluation in the Data Structures and Algorithms course, this study presents DSA-RubricEval, a pedagogically grounded and preliminarily reliability-assessed dataset. The dataset constitutes the primary contribution of this work, providing a structured benchmark aligned with Bloom’s taxonomy and validated through multi-rater Interclass Correlation Coefficient ICC analysis. The dataset consists of twelve expert-designed, Bloom-aligned, questions scored on multiple rubric dimensions using an ordinal scale. Five independent evaluators scored student responses, enabling rigorous validation of human judgement based on the Intraclass Correlation Coefficient (ICC) analysis. The results indicate good to excellent average-measure reliability across most rubric dimensions, justifying the use of aggregated human scores as aggregated reference labels for automated assessment. Automated scoring was explored as a proof of concept to demonstrate the applicability of the dataset for AI-assisted assessment and formulated as an ordinal, rubric-level prediction task and evaluated using pedagogically motivated agreement measures, showing high tolerance-based agreement with human judges. The proposed research identifies sources of assessor subjectivity and explores methods to mitigate them, while reducing the workload of grading as well as correlate with learning outcomes. Moreover, the proposed technique supports lower-order cognitive skills but is not best suited for higher-order cognitive tasks. Overall, this study introduces a preliminarily reliability-assessed, rubric-based benchmark dataset intended to support exploratory research on pedagogically meaningful AI-assisted assessment of open-ended technical responses.
J. K. Sheikh, Hemant Kumar Soni· Discover Education· 0 citations