Skip to content

Papers

Selected publications and preprints from the AIXpert project. Each entry links to arXiv where available.


Harness-Aware Evaluation of LLM Agents: Systems, Benchmarks, and Protocols

Paper (Preprint) · Preprint Code · GitHub Project · Project

Authors: Ahmed Y. Radwan, Athanasios V. Vasilakos, Shaina Raza.

Survey of agent evaluation that separates backbone-model effects from the harness, environment, and scoring setup.


A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems

Paper (Preprint) · arXiv

Authors: Shaina Raza, Ahmed Y. Radwan, Imran Liaquat, Kathryn Hume.

Framework for evaluating trustworthy AI across LLMs, agents, and multimodal systems. It uses eight trustworthiness dimensions, shared scoring bands, safety vetoes, and evaluation-quality checks, with links to the EU AI Act, ISO standards, and the NIST AI RMF.


Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions

Paper (NeurIPS 2026 TAE) · arXiv Code · GitHub Project · Project

Authors: Ahmed El Kady, Aravind Narayanan, Yani Ioannou, Shaina Raza.

Study of how batching, quantization, and reduced benchmark subsets change responsible-AI conclusions on BBQ and BBQ-V: accuracy can hold while bias, reasoning quality, and subgroup results shift, and lower-precision inference can use more GPU energy rather than less.


FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making

Paper (Preprint) · arXiv Code · GitHub Project · Project Dataset · Hugging Face

Authors: Vahid Reza Khazaie, Ahmed Y. Radwan, Shaina Raza.

Fairness benchmark for vision-language models on high-stakes hiring, legal, and healthcare decisions from a face photograph. Reports parity, soundness, demographic association, and free-text bias.


AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA

Paper (NeurIPS 2026 AABA4ET) · arXiv Code · GitHub Project · Project

Authors: Aravind Narayanan, Shaina Raza.

Multi-agent pipeline for auditable, on-premise financial chart QA. Queries are split into planning, OCR, legend grounding, visual inspection, and verification, and every step is recorded in a Model Evaluation Packet (MEP). +7.68 pp over a zero-shot baseline on FinMME, with full open-weights deployment for data residency.


Detecting Deception, Not Deepfakes: Why Media Forensics Needs Social Theories

Paper (NeurIPS 2026 TAE) · arXiv Project · Project

Authors: Jessee Ho, Shweta Khushu, Shaina Raza.

Position paper arguing that artifact-based deepfake detectors miss interactive deception, and that forensics also needs a communication layer covering speech acts, conversation, and influence.


Deepfakes in the Foundation-Model Era: A Survey of Forensics, Generation and Distribution Across Social Media Lifecycle

Paper (Preprint) · Preprint Code · GitHub Project · Project

Authors: Shaina Raza, Jessee Ho, Ahmed Y. Radwan, Mohamed Hafez.

Survey of deepfakes from text-to-video and jointly generated audio-video models through how platforms compress, repost, and spread them, and of what forensic evidence is left when a detection decision has to be made.


Position: Sustainable Open-Source AI Requires Tracking the Cumulative Footprint of Derivatives

Paper (ICML 2026) · ICML Code · GitHub Project · Project PyPI · PyPI Dashboard · Dashboard

Authors: Shaina Raza, Iuliia Zarubiieva, Ahmed Radwan, Nathaniel Lesperance, Deval Pandya, Sedef Akinli Kocak, Graham Taylor.

Position paper proposing Data and Impact Accounting (DIA), now also an open-source toolkit, to track energy, water, and CO₂ impact of open-source AI models and their derivatives across lineages.


UnBias-Plus: Detect, Explain, and Rewrite Bias

Paper · arXiv Code · GitHub Project · Project Demo · Demo

Authors: Ahmed Y. Radwan, Ahmed ElKady, Sindhuja Chaduvula, Mohamed Hafez, Amrit Krishnan, Shaina Raza.

Open-source toolkit unifying segment-level bias classification, biased span localization, neutral text rewriting, and per-decision reasoning. Available via Python, CLI, REST API, and web interfaces.


Detecting and Reasoning About Bias in Multimodal Content

Paper (IASEAI 2026) · IASEAI

Authors: Shaina Raza et al.

Work on detecting and reasoning about bias in multimodal content.


DiagFlowBench: Evaluating How Language Models Handle Off-Procedure Inputs in Grounded Diagnostic Dialogue

Paper (EMNLP 2026 Industry) · arXiv

Authors: Guillermo Gil de Avalle, Laura Maruster, Shaina Raza, Christos Emmanouilidis.

Benchmark evaluating how language models handle off-procedure inputs in grounded diagnostic dialogue.


HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation

Paper (ACM TIST) · ACM Project · Project Dataset · Hugging Face

Authors: Shaina Raza, Aravind Narayanan, Vahid Reza Khazaie, Ashmal Vayani, Ahmed Y. Radwan, Mukund S. Chettiar, Amandeep Singh, Mubarak Shah, Deval Pandya.

Fairness-focused vision-language benchmark evaluating large multimodal models across human-centric demographics and HCAI principles.


SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

Paper (EMNLP 2026 Main) · arXiv Code · GitHub Dataset · Hugging Face Leaderboard · Leaderboard

Authors: Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum, Deval Pandya, Shaina Raza.

SONIC-O1, a fully human-verified real-world audio-video benchmark with 4,958 annotations across 13 conversational domains. We evaluate multimodal models on video summarization, evidence-grounded QA, and temporal event localization, and release an extensible evaluation suite to support reproducible benchmarking and robustness analysis.


Evaluating and Regulating Agentic AI: A Study of Benchmarks, Metrics and Regulation

Paper (Information Fusion, Elsevier 2026) · ScienceDirect Code · GitHub Project · Project

Authors: Azib Farooq, Shaina Raza, Nazmul Karim, Hasan Iqbal, Athanasios V. Vasilakos, Christos Emmanouilidis.

Survey of benchmarks, metrics, and governance for evaluating agentic AI in single- and multi-agent systems, toward trustworthy and auditable agents.


Generative Deepfake Videos in the Foundation-Model Era: A Timeline of Eroding Trust in Visual Evidence

Paper (MAD '26, ACM) · ACM

Authors: Shaina Raza, Jessee Ho, Mahveen Raza, Christos Emmanouilidis.

Review of the post-artifact deepfake era using a forensic-assumption framework (physiological integrity, temporal coherence, geometric consistency, semantic consistency, provenance) and five detection paradigms (spatial-frequency, temporal, multimodal, vision-language, and agentic), plus an audit of thirteen benchmarks.


TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems

Paper (AI Open, Elsevier 2026) · AI Open

Authors: Shaina Raza, Ranjan Sapkota, Manoj Karkee, Christos Emmanouilidis.

A review of trust, risk, and security management (TRiSM) in LLM-based agentic and multi-agent systems.


Just as Humans Need Vaccines, So Do Models: Model Immunization to Combat Falsehoods

Paper (WCCI 2026, IJCNN) · arXiv Code · GitHub Project · Project

Authors: Shaina Raza, Rizwan Qureshi, Azib Farooq, Marcelo Lotif, Aman Chadha, Deval Pandya, Christos Emmanouilidis.

Position paper on model immunization: SFT with small doses of (false claim, correction) pairs alongside truthful data to supervise falsehoods directly. Across four open-weight families, +12 TruthfulQA and +30 misinformation-rejection points with negligible capability loss.


From Features to Actions: Explainability in Traditional and Agentic AI

Paper · arXiv Code · GitHub Project · Project

Authors: Sindhuja Chaduvula, Jessee Ho, Kina Kim, Aravind Narayanan, Mahshid Alinoori, Muskan Garg, Dhanesh Ramachandram, Shaina Raza.

Compares attribution-based explanations (SHAP, LIME) with trace-based diagnostics across static and agentic settings. Attribution is stable for static prediction (Spearman ρ = 0.86) but fails to diagnose agentic failures; trace-grounded rubrics localize breakdowns (e.g. state-tracking inconsistency 2.7× more in failed runs, −49% success), motivating trajectory-level explainability for agentic systems.


Reducing Hallucinations in LLMs via Factuality-Aware Preference Learning

Paper (ACL 2026 Findings) · ACL Anthology Code · GitHub Dataset · Hugging Face Project · Project

Authors: Sindhuja Chaduvula, Ahmed Y. Radwan, Azib Farooq, Yani Ioannou, Shaina Raza.

Preference-learning method (F-DPO) that targets factuality directly, improving factuality scores while reducing hallucination rates across multiple open-weight LLMs.


Transparency in Agentic AI: A Survey of Interpretability, Explainability, and Governance

Paper · arXiv Project · Project

Authors: Shaina Raza, Ahmed Y. Radwan, Sindhuja Chaduvula, Mahshid Alinoori, Christos Emmanouilidis.

Survey of interpretability and explainability for LLM-based agents with planning, memory, and tool use. Organized by what is inspected, what is assured, the methods used, the lifecycle stage, and who the explanation is for.


Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment

Paper (NeurIPS 2025 LLM-eval Workshop) · arXiv Code · GitHub

Authors: Aravind Narayanan, Vahid Reza Khazaie, Shaina Raza.

Benchmarking vision-language models with social-cue news images and LLM-as-judge assessment.


Responsible Agentic Reasoning and AI Agents—A Critical Survey

Paper · arXiv

Authors: Shaina Raza (Vector Institute), Ranjan Sapkota, Manoj Karkee (Cornell University), Christos Emmanouilidis (University of Groningen).

Critical survey of responsible agentic reasoning and AI agents.