Abstract
Evaluating Artificial Intelligence (AI) and data science models is crucial to ensure their reliability, fairness, and applicability in real-world scenarios. This paper highlights best practices for model evaluation, emphasizing the importance of selecting appropriate metrics aligned with business or research goals. Key considerations include using robust validation strategies (e.g., cross-validation), monitoring for overfitting, and ensuring data splits preserve class distributions. Fairness, interpretability, and reproducibility are essential, particularly in high-stakes domains like healthcare or finance. Additionally, evaluating models across multiple datasets or demographic subgroups helps uncover biases and improve generalizability. Adopting standardized reporting practices and
open-source benchmarks further strengthens the evaluation process. By adhering to these practices, practitioners can build more trustworthy and effective AI systems.
open-source benchmarks further strengthens the evaluation process. By adhering to these practices, practitioners can build more trustworthy and effective AI systems.
| Original language | English |
|---|---|
| Title of host publication | INFORMATIK 2025 : The Wide Open - Offenheit von Source bis Science, 16.-19.September 2025 Potsdam |
| Editors | Ulrike Lucke, Stefan Stieglitz, Falk Uebernickel, Anna-Lena Lamprecht, Maike Klein |
| Number of pages | 9 |
| Place of Publication | Bonn |
| Publisher | Gesellschaft für Informatik e.V. |
| Publication date | 2025 |
| Pages | 1211-1219 |
| DOIs | |
| Publication status | Published - 2025 |
Bibliographical note
Publisher Copyright:© 2025 Gesellschaft fur Informatik (GI). All rights reserved.
Research areas and keywords
- AI
- Best Practices
- Data Science
- Evaluation
- Machine Learning
ASJC Scopus Subject Areas
- Computer Science Applications
Fingerprint
Dive into the research topics of 'Best Practices in AI and Data Science Models Evaluation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver