Papers¶
Selected publications and preprints from the AIXpert project. Each entry links to arXiv where available.
Harness-Aware Evaluation of LLM Agents: Systems, Benchmarks, and Protocols¶
Paper (Preprint) · Code ·
Project ·
Authors: Ahmed Y. Radwan, Athanasios V. Vasilakos, Shaina Raza.
Survey of agent evaluation that separates backbone-model effects from the harness, environment, and scoring setup.
A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems¶
Authors: Shaina Raza, Ahmed Y. Radwan, Imran Liaquat, Kathryn Hume.
Framework for evaluating trustworthy AI across LLMs, agents, and multimodal systems. It uses eight trustworthiness dimensions, shared scoring bands, safety vetoes, and evaluation-quality checks, with links to the EU AI Act, ISO standards, and the NIST AI RMF.
Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions¶
Paper (NeurIPS 2026 TAE) · Code ·
Project ·
Authors: Ahmed El Kady, Aravind Narayanan, Yani Ioannou, Shaina Raza.
Study of how batching, quantization, and reduced benchmark subsets change responsible-AI conclusions on BBQ and BBQ-V: accuracy can hold while bias, reasoning quality, and subgroup results shift, and lower-precision inference can use more GPU energy rather than less.
FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making¶
Paper (Preprint) · Code ·
Project ·
Dataset ·
Authors: Vahid Reza Khazaie, Ahmed Y. Radwan, Shaina Raza.
Fairness benchmark for vision-language models on high-stakes hiring, legal, and healthcare decisions from a face photograph. Reports parity, soundness, demographic association, and free-text bias.
AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA¶
Paper (NeurIPS 2026 AABA4ET) · Code ·
Project ·
Authors: Aravind Narayanan, Shaina Raza.
Multi-agent pipeline for auditable, on-premise financial chart QA. Queries are split into planning, OCR, legend grounding, visual inspection, and verification, and every step is recorded in a Model Evaluation Packet (MEP). +7.68 pp over a zero-shot baseline on FinMME, with full open-weights deployment for data residency.
Detecting Deception, Not Deepfakes: Why Media Forensics Needs Social Theories¶
Paper (NeurIPS 2026 TAE) · Project ·
Authors: Jessee Ho, Shweta Khushu, Shaina Raza.
Position paper arguing that artifact-based deepfake detectors miss interactive deception, and that forensics also needs a communication layer covering speech acts, conversation, and influence.
Deepfakes in the Foundation-Model Era: A Survey of Forensics, Generation and Distribution Across Social Media Lifecycle¶
Paper (Preprint) · Code ·
Project ·
Authors: Shaina Raza, Jessee Ho, Ahmed Y. Radwan, Mohamed Hafez.
Survey of deepfakes from text-to-video and jointly generated audio-video models through how platforms compress, repost, and spread them, and of what forensic evidence is left when a detection decision has to be made.
Position: Sustainable Open-Source AI Requires Tracking the Cumulative Footprint of Derivatives¶
Paper (ICML 2026) · Code ·
Project ·
PyPI ·
Dashboard ·
Authors: Shaina Raza, Iuliia Zarubiieva, Ahmed Radwan, Nathaniel Lesperance, Deval Pandya, Sedef Akinli Kocak, Graham Taylor.
Position paper proposing Data and Impact Accounting (DIA), now also an open-source toolkit, to track energy, water, and CO₂ impact of open-source AI models and their derivatives across lineages.
UnBias-Plus: Detect, Explain, and Rewrite Bias¶
Paper · Code ·
Project ·
Demo ·
Authors: Ahmed Y. Radwan, Ahmed ElKady, Sindhuja Chaduvula, Mohamed Hafez, Amrit Krishnan, Shaina Raza.
Open-source toolkit unifying segment-level bias classification, biased span localization, neutral text rewriting, and per-decision reasoning. Available via Python, CLI, REST API, and web interfaces.
Detecting and Reasoning About Bias in Multimodal Content¶
Authors: Shaina Raza et al.
Work on detecting and reasoning about bias in multimodal content.
DiagFlowBench: Evaluating How Language Models Handle Off-Procedure Inputs in Grounded Diagnostic Dialogue¶
Authors: Guillermo Gil de Avalle, Laura Maruster, Shaina Raza, Christos Emmanouilidis.
Benchmark evaluating how language models handle off-procedure inputs in grounded diagnostic dialogue.
HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation¶
Paper (ACM TIST) · Project ·
Dataset ·
Authors: Shaina Raza, Aravind Narayanan, Vahid Reza Khazaie, Ashmal Vayani, Ahmed Y. Radwan, Mukund S. Chettiar, Amandeep Singh, Mubarak Shah, Deval Pandya.
Fairness-focused vision-language benchmark evaluating large multimodal models across human-centric demographics and HCAI principles.
SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding¶
Paper (EMNLP 2026 Main) · Code ·
Dataset ·
Leaderboard ·
Authors: Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum, Deval Pandya, Shaina Raza.
SONIC-O1, a fully human-verified real-world audio-video benchmark with 4,958 annotations across 13 conversational domains. We evaluate multimodal models on video summarization, evidence-grounded QA, and temporal event localization, and release an extensible evaluation suite to support reproducible benchmarking and robustness analysis.
Evaluating and Regulating Agentic AI: A Study of Benchmarks, Metrics and Regulation¶
Paper (Information Fusion, Elsevier 2026) · Code ·
Project ·
Authors: Azib Farooq, Shaina Raza, Nazmul Karim, Hasan Iqbal, Athanasios V. Vasilakos, Christos Emmanouilidis.
Survey of benchmarks, metrics, and governance for evaluating agentic AI in single- and multi-agent systems, toward trustworthy and auditable agents.
Generative Deepfake Videos in the Foundation-Model Era: A Timeline of Eroding Trust in Visual Evidence¶
Authors: Shaina Raza, Jessee Ho, Mahveen Raza, Christos Emmanouilidis.
Review of the post-artifact deepfake era using a forensic-assumption framework (physiological integrity, temporal coherence, geometric consistency, semantic consistency, provenance) and five detection paradigms (spatial-frequency, temporal, multimodal, vision-language, and agentic), plus an audit of thirteen benchmarks.
TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems¶
Paper (AI Open, Elsevier 2026) ·
Authors: Shaina Raza, Ranjan Sapkota, Manoj Karkee, Christos Emmanouilidis.
A review of trust, risk, and security management (TRiSM) in LLM-based agentic and multi-agent systems.
Just as Humans Need Vaccines, So Do Models: Model Immunization to Combat Falsehoods¶
Paper (WCCI 2026, IJCNN) · Code ·
Project ·
Authors: Shaina Raza, Rizwan Qureshi, Azib Farooq, Marcelo Lotif, Aman Chadha, Deval Pandya, Christos Emmanouilidis.
Position paper on model immunization: SFT with small doses of (false claim, correction) pairs alongside truthful data to supervise falsehoods directly. Across four open-weight families, +12 TruthfulQA and +30 misinformation-rejection points with negligible capability loss.
From Features to Actions: Explainability in Traditional and Agentic AI¶
Authors: Sindhuja Chaduvula, Jessee Ho, Kina Kim, Aravind Narayanan, Mahshid Alinoori, Muskan Garg, Dhanesh Ramachandram, Shaina Raza.
Compares attribution-based explanations (SHAP, LIME) with trace-based diagnostics across static and agentic settings. Attribution is stable for static prediction (Spearman ρ = 0.86) but fails to diagnose agentic failures; trace-grounded rubrics localize breakdowns (e.g. state-tracking inconsistency 2.7× more in failed runs, −49% success), motivating trajectory-level explainability for agentic systems.
Reducing Hallucinations in LLMs via Factuality-Aware Preference Learning¶
Paper (ACL 2026 Findings) · Code ·
Dataset ·
Project ·
Authors: Sindhuja Chaduvula, Ahmed Y. Radwan, Azib Farooq, Yani Ioannou, Shaina Raza.
Preference-learning method (F-DPO) that targets factuality directly, improving factuality scores while reducing hallucination rates across multiple open-weight LLMs.
Transparency in Agentic AI: A Survey of Interpretability, Explainability, and Governance¶
Authors: Shaina Raza, Ahmed Y. Radwan, Sindhuja Chaduvula, Mahshid Alinoori, Christos Emmanouilidis.
Survey of interpretability and explainability for LLM-based agents with planning, memory, and tool use. Organized by what is inspected, what is assured, the methods used, the lifecycle stage, and who the explanation is for.
Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment¶
Paper (NeurIPS 2025 LLM-eval Workshop) · Code ·
Authors: Aravind Narayanan, Vahid Reza Khazaie, Shaina Raza.
Benchmarking vision-language models with social-cue news images and LLM-as-judge assessment.
Responsible Agentic Reasoning and AI Agents—A Critical Survey¶
Authors: Shaina Raza (Vector Institute), Ranjan Sapkota, Manoj Karkee (Cornell University), Christos Emmanouilidis (University of Groningen).
Critical survey of responsible agentic reasoning and AI agents.