Updates¶
Full list of recent papers, releases, and news.
-
SONIC-O1 accepted at EMNLP 2026 (Main). Real-world audio-video benchmark for evaluating multimodal LLMs. Project page · Preprint · Code · Dataset · Leaderboard.
-
DiagFlowBench accepted at EMNLP 2026 (Industry). Benchmark evaluating how language models handle off-procedure inputs in grounded diagnostic dialogue. Preprint.
-
AgentFinVQA accepted at NeurIPS 2026 AABA4ET. Auditable financial chart QA pipeline for regulated, on-premise use. Project page · Preprint · Code.
-
Detecting Deception, Not Deepfakes accepted at NeurIPS 2026 TAE. Position paper on why artifact detectors miss interactive deception, adding a communication layer for speech acts, conversation, and influence. Project page · Preprint.
-
Green AI Workshop (ECML-PKDD 2026). Shaina Raza delivered a keynote presenting DIA and the ICML 2026 Spotlight Position Paper on tracking the cumulative energy, carbon, and water footprint of model derivatives. Program · Project.
-
Harness-Aware Evaluation of LLM Agents. Survey of how memory, tools, and execution software around an LLM change reported agent scores. Project page · Preprint · Code.
-
Unified evaluation framework. Framework for trustworthy evaluation across LLMs, agentic AI, and multimodal systems, with shared scoring bands and links to the EU AI Act, ISO, and NIST AI RMF. Preprint.
-
Stress-Testing Efficient Responsible-AI Evaluation accepted at NeurIPS 2026 TAE. Study of how batching, quantization, and smaller subsets change bias-benchmark conclusions and energy use on BBQ and BBQ-V. Project page · Preprint · Code.
-
FairLens. Fairness benchmark for VLMs on high-stakes hiring, legal, and healthcare decisions from a face photograph, measuring parity, soundness, and refusal to judge from appearance. Project page · Preprint · Code · Dataset.
-
Deepfakes in the Foundation-Model Era. Survey of forensics, generation, and platform distribution. Detectors built for older deepfakes often miss newer ones circulating online. Project page · Preprint · Code.
-
UnBias-Plus (Featured Project). Toolkit for bias detection, explanation, severity assessment, and debiasing in text (segmentation, severity scoring, reasoning, replacements, and full rewrite). Paper · Project page · Code · Demo · PyPI
-
DIA (Data & Impact Accounting) (Featured Project). Open-source toolkit that tracks the energy, water, and CO₂ footprint of open-source AI models and their derivatives, now released as the
ai-impact-accountingpackage with code, docs, and an interactive dashboard. Project page · Code · PyPI · Dashboard · ICML 2026 paper -
HumaniBench accepted at ACM TIST. Fairness-focused vision-language benchmark published in ACM Transactions on Intelligent Systems and Technology. Project page · Paper.
-
DIA in the press. Nosian Magazine's AI's Unasked Question: What Did That Cost? features DIA and the team behind it (Amrit Krishnan, Shaina Raza, PhD, and Ahmed Y. Radwan; written by Sally Work), on why so few ICML papers measure AI's environmental cost and how the
ai-impact-accountingpackage helps. -
UnBias-Plus in the press. Independent coverage of the open-source writing-bias detector and rewrite toolkit (healthcare, HR/AI workflows, and AI training-data screening): CAN Health (8 Jul 2026) · ChannelLife (2 Jul 2026) · AI Loop (22 Jul 2026) · BornCity (3 Jul 2026) · TechTalent.ca (Jul 2026) · GlobeNewswire (30 Jun 2026) · BNN Bloomberg (30 Jun 2026).
-
The AmberMac Show (Ep. 076). Kathryn Hume, VP of AI Engineering at Vector Institute, discusses AI bias and how AI is reshaping high-stakes decisions like hiring and mortgage approvals, and what's needed to reduce bias and discrimination.
-
MoE blog. Mixture of Experts: From Sparse Routing to Multimodal Deployment.
-
Detecting and Reasoning About Bias in Multimodal Content. IASEAI 2026. Paper.
-
DiagFlowBench. Preprint: Evaluating How Language Models Handle Off-Procedure Inputs in Grounded Diagnostic Dialogue.
-
AIXpert General Assembly: Barcelona 2026. The AIXPERT consortium met at the Barcelona Supercomputing Center (3-4 June 2026) for its second-year General Assembly and Technical Meeting. Partners aligned the technical roadmap, validated implementation plans across five pilot domains (healthcare, recruitment, manufacturing, education, and creative industries), and prepared the next phase of platform development. Read more.
- Toronto Machine Learning Summit. Ahmed Y. Radwan presented SONIC-O1 at the Toronto Machine Learning Summit (16-19 June 2026, Toronto). SONIC-O1 is a real-world benchmark for evaluating multimodal LLMs on audio-video understanding across 13 conversational domains, featuring a public leaderboard for model comparison and fairness analysis.
Project page · Code · Dataset · Leaderboard
- HAICON26 & Vector-Helmholtz Munich MOU. Shaina Raza, PhD presented at the Helmholtz AI Conference 2026: AI for Science (8-11 June 2026, Munich). The visit coincided with the signing of a Memorandum of Understanding (MOU) between Vector Institute and Helmholtz Munich's Computational Health Center, establishing a formal framework for joint research, researcher exchanges, faculty affiliations, and coordinated Germany-Canada funding programs.
Helmholtz Munich · Press release (GlobeNewsWire) · AI Magazine · BNN Bloomberg
-
ICML 2026 Poster. Position: Sustainable Open-Source AI Requires Tracking the Cumulative Footprint of Derivatives accepted as a poster at ICML 2026. Shaina Raza, Ahmed Radwan, and co-authors propose Data and Impact Accounting (DIA) to make the cumulative environmental footprint of open-source AI model derivatives visible and accountable. Project page · Code · PyPI.
-
Deepfake Video Review. Generative Deepfake Videos in the Foundation-Model Era: A Timeline of Eroding Trust in Visual Evidence published at MAD '26 (ACM Workshop on Multimedia AI against Disinformation, June 2026). Shaina Raza, Jessee Ho, Mahveen Raza, Christos Emmanouilidis.
-
The Peak Emerging Leaders 2026. Shaina Raza, PhD has been recognized in The Peak's Emerging Leaders 2026 list, celebrating Canada's most promising young leaders in Artificial Intelligence.
-
FairSense-AgentiX. Agentic fairness and AI-risk platform for text, images, and datasets (FastAPI, WebSocket, React UI, Python API). Project page · Code · PyPI.
-
HAICON26. Shaina Raza, PhD presenting at the Helmholtz AI Conference 2026: AI for Science (8-11 June 2026, Munich, Germany).
-
Toronto Machine Learning Summit. Ahmed Y. Radwan presenting SONIC-O1 at the Toronto Machine Learning Summit (16-19 June 2026). Project page · Code · Dataset · Leaderboard.
-
AI4Good Lab 2026. Shaina Raza, PhD and Ahmed Y. Radwan presented UnBias-Plus and research on disinformation and misinformation detection to the AI4Good Lab 2026 cohort (May-June 2026, Toronto). The AI4Good Lab is a full-time summer ML training program for women and gender diverse people across Canada, hosted in Toronto in partnership with Vector Institute and CIFAR.
-
Evaluating and Regulating Agentic AI. Published in Information Fusion. Evaluating and Regulating Agentic AI: A Study of Benchmarks, Metrics, and Regulation. Project page · Code.
-
EU cluster webinar: AI in public services. We contributed to the webinar AI-Enabled Public Services: Building Resilience and Accountability (20 April 2026, 10:00 CEST). The session was the second joint online event organized by the EU-funded projects TANGO, AI4REALNET, HumAIne, THEMIS 5.0, and Peer AI, bringing together speakers on how these initiatives use AI in public-service settings, with emphasis on resilience and accountability. Registration was open to the general public, researchers, and policy makers; registered participants received the detailed agenda in the weeks before the event.
-
Model immunization (AI vaccine). Accepted at WCCI 2026 (IJCNN). Just as Humans Need Vaccines, So Do Models: Model Immunization to Combat Falsehoods (arXiv). Project page · Code.
-
F-DPO. ACL 2026 Findings. Reducing Hallucinations in LLMs via Factuality-Aware Preference Learning. Project page · Code · Dataset.
-
TRiSM for Agentic AI accepted. Paper accepted at AI Open, Elsevier 2026. A review of trust, risk, and security management in LLM-based agentic multi-agent systems.
-
AgentFinVQA. Multi-agent pipeline for auditable, on-premise financial chart QA with traceable Model Evaluation Packets. Paper · Project page · Code.
-
Remarkable 2026. We presented AIXpert projects at Remarkable 2026.
- SONIC-O1 Multi-Agent. We released a compound multi-agent system for audio-video understanding. Code.
- From Features to Actions. Paper: Explainability in Traditional and Agentic AI Systems (arXiv). Code and project page.
- Transparency in Agentic AI. Survey: A Survey of Interpretability, Explainability, and Governance (arXiv). Project page.
- AIXpert news. Our work was highlighted on the AIXpert project website: Advancing Trustworthy, Explainable, and Responsible AI at NeurIPS 2025 (Bias in the Picture, HumaniBench, Carbon Literacy, and more).
- SONIC-O1. Paper: A Real-World Benchmark for Evaluating MLLMs on Audio-Video Understanding (arXiv).
- SONIC-O1. Dataset on Hugging Face (231 videos, ~60h, 4,958 QAs, 13 domains, demographic metadata).
- SONIC-O1. Code and evaluation pipeline (summarization, MCQ, temporal localization).
- SONIC-O1. Leaderboard for model comparisons and fairness analysis.
- F-DPO. Factuality-aware preference dataset on Hugging Face.
- F-DPO. Code and project page (factuality-aware DPO, no reward model).
- Released data generation pipeline (multimodal, configurable, agent-orchestrated).
- Single-agent pipeline prototype for rapid dataset bootstrapping.
- Paper: Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment (NeurIPS 2025 LLM-eval Workshop)
- Paper: TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems
- Paper: Responsible Agentic Reasoning and AI Agents—A Critical Survey (arXiv)
- Paper: Single-Agent TRiSM (Poster, NeurIPS LAW)