Skip to content

Updates

Full list of recent papers, releases, and news.


  • SONIC-O1 accepted at EMNLP 2026 (Main). Real-world audio-video benchmark for evaluating multimodal LLMs. Project page · Preprint · Code · Dataset · Leaderboard.

  • DiagFlowBench accepted at EMNLP 2026 (Industry). Benchmark evaluating how language models handle off-procedure inputs in grounded diagnostic dialogue. Preprint.

  • AgentFinVQA accepted at NeurIPS 2026 AABA4ET. Auditable financial chart QA pipeline for regulated, on-premise use. Project page · Preprint · Code.

  • Detecting Deception, Not Deepfakes accepted at NeurIPS 2026 TAE. Position paper on why artifact detectors miss interactive deception, adding a communication layer for speech acts, conversation, and influence. Project page · Preprint.

  • Green AI Workshop (ECML-PKDD 2026). Shaina Raza delivered a keynote presenting DIA and the ICML 2026 Spotlight Position Paper on tracking the cumulative energy, carbon, and water footprint of model derivatives. Program · Project.

  • Harness-Aware Evaluation of LLM Agents. Survey of how memory, tools, and execution software around an LLM change reported agent scores. Project page · Preprint · Code.

  • Unified evaluation framework. Framework for trustworthy evaluation across LLMs, agentic AI, and multimodal systems, with shared scoring bands and links to the EU AI Act, ISO, and NIST AI RMF. Preprint.

  • Stress-Testing Efficient Responsible-AI Evaluation accepted at NeurIPS 2026 TAE. Study of how batching, quantization, and smaller subsets change bias-benchmark conclusions and energy use on BBQ and BBQ-V. Project page · Preprint · Code.

  • FairLens. Fairness benchmark for VLMs on high-stakes hiring, legal, and healthcare decisions from a face photograph, measuring parity, soundness, and refusal to judge from appearance. Project page · Preprint · Code · Dataset.

  • Deepfakes in the Foundation-Model Era. Survey of forensics, generation, and platform distribution. Detectors built for older deepfakes often miss newer ones circulating online. Project page · Preprint · Code.

  • UnBias-Plus (Featured Project). Toolkit for bias detection, explanation, severity assessment, and debiasing in text (segmentation, severity scoring, reasoning, replacements, and full rewrite). Paper · Project page · Code · Demo · PyPI

  • DIA (Data & Impact Accounting) (Featured Project). Open-source toolkit that tracks the energy, water, and CO₂ footprint of open-source AI models and their derivatives, now released as the ai-impact-accounting package with code, docs, and an interactive dashboard. Project page · Code · PyPI · Dashboard · ICML 2026 paper

  • HumaniBench accepted at ACM TIST. Fairness-focused vision-language benchmark published in ACM Transactions on Intelligent Systems and Technology. Project page · Paper.

  • DIA in the press. Nosian Magazine's AI's Unasked Question: What Did That Cost? features DIA and the team behind it (Amrit Krishnan, Shaina Raza, PhD, and Ahmed Y. Radwan; written by Sally Work), on why so few ICML papers measure AI's environmental cost and how the ai-impact-accounting package helps.

  • UnBias-Plus in the press. Independent coverage of the open-source writing-bias detector and rewrite toolkit (healthcare, HR/AI workflows, and AI training-data screening): CAN Health (8 Jul 2026) · ChannelLife (2 Jul 2026) · AI Loop (22 Jul 2026) · BornCity (3 Jul 2026) · TechTalent.ca (Jul 2026) · GlobeNewswire (30 Jun 2026) · BNN Bloomberg (30 Jun 2026).

  • The AmberMac Show (Ep. 076). Kathryn Hume, VP of AI Engineering at Vector Institute, discusses AI bias and how AI is reshaping high-stakes decisions like hiring and mortgage approvals, and what's needed to reduce bias and discrimination.

  • MoE blog. Mixture of Experts: From Sparse Routing to Multimodal Deployment.

  • Detecting and Reasoning About Bias in Multimodal Content. IASEAI 2026. Paper.

  • DiagFlowBench. Preprint: Evaluating How Language Models Handle Off-Procedure Inputs in Grounded Diagnostic Dialogue.

  • AIXpert General Assembly: Barcelona 2026. The AIXPERT consortium met at the Barcelona Supercomputing Center (3-4 June 2026) for its second-year General Assembly and Technical Meeting. Partners aligned the technical roadmap, validated implementation plans across five pilot domains (healthcare, recruitment, manufacturing, education, and creative industries), and prepared the next phase of platform development. Read more.

  • Toronto Machine Learning Summit. Ahmed Y. Radwan presented SONIC-O1 at the Toronto Machine Learning Summit (16-19 June 2026, Toronto). SONIC-O1 is a real-world benchmark for evaluating multimodal LLMs on audio-video understanding across 13 conversational domains, featuring a public leaderboard for model comparison and fairness analysis.

Project page · Code · Dataset · Leaderboard

  • HAICON26 & Vector-Helmholtz Munich MOU. Shaina Raza, PhD presented at the Helmholtz AI Conference 2026: AI for Science (8-11 June 2026, Munich). The visit coincided with the signing of a Memorandum of Understanding (MOU) between Vector Institute and Helmholtz Munich's Computational Health Center, establishing a formal framework for joint research, researcher exchanges, faculty affiliations, and coordinated Germany-Canada funding programs.

Helmholtz Munich · Press release (GlobeNewsWire) · AI Magazine · BNN Bloomberg