International Journal of Technology and Applied Science

E-ISSN: 2230-9004     Impact Factor: 9.914

A Widely Indexed Open Access Peer Reviewed Multidisciplinary Bi-monthly Scholarly International Journal

Call for Paper Volume 17 Issue 8 (August 2026) Submit your research before the last 3 days of this month to publish your research paper in the current issue.

Measuring Trust in Generative AI: A Framework for Accuracy, Bias, and Safety Metrics

Author(s) Mr. Rajesh Lingam, Ms. Meena Masoodi
Country United States
Abstract Trust in Generative AI systems depends on demonstrable, measurable evidence that these systems behave accurately, fairly, and safely across diverse user populations and use cases. Despite growing industry and regulatory interest in trustworthy AI, practical frameworks for operationalizing trust measurement in production systems remain underdeveloped. Existing benchmarks such as HELM and BIG-Bench evaluate capability breadth but do not produce actionable trust scores suitable for continuous production monitoring. This paper proposes the Trust Measurement Framework (TMF) for Generative AI, defining quantifiable metrics across three dimensions: accuracy (semantic correctness, factual grounding, and citation precision), bias (demographic parity delta, equalized odds difference, and counterfactual sensitivity), and safety (harmful content rate, boundary robustness index, and adversarial resistance score). We describe metric definitions, measurement methodologies, and a weighted aggregation strategy that produces a single composite Trust Score for production LLM systems. Validated across three successive model versions in an enterprise document intelligence platform serving over ten million users, the TMF demonstrated stable predictive validity for user satisfaction outcomes (Pearson r=0.81 across three model versions) and correctly detected a safety regression in a controlled pre-deployment scenario that accuracy-only evaluation would have missed. We discuss metric selection trade-offs, calibration approaches, and integration patterns with automated evaluation pipelines. The TMF provides a practical, reproducible foundation for organizations that must demonstrate trustworthy AI behavior to regulators, customers, and internal governance bodies.
Keywords Trustworthy AI, generative AI evaluation, large language models, bias measurement, AI safety, trust scoring, enterprise AI, responsible AI.
Field Computer > Artificial Intelligence / Simulation / Virtual Reality
Published In Volume 14, Issue 9, September 2023
Published On 2023-09-19
DOI https://doi.org/10.71097/IJTAS.v14.i9.1349

Share this