Skip to content

Vector AIXpert: Responsible AI Infrastructure for Fairness, Explainability, and Evaluation

Vector Institute's contribution to the AIXpert Project: tools, benchmarks, and research for explainable, accountable, and fair AI.

This project represents the Vector Institute's research contributions to the AIXpert Horizon Europe initiative. It focuses on developing tools, datasets, and evaluation pipelines for fairness-aware generative AI and explainable AI systems.


What we do

Vector's contribution to AIXpert spans four core areas:

  • Explainable & accountable AI. Tools and benchmarks for interpretability, fairness, and transparency in generative and multimodal AI.
  • Trustworthy agentic AI. Transparent, auditable, human-in-the-loop agentic systems with measurable trustworthiness metrics.
  • Multimodal evaluation. Benchmarks and datasets for audio-video understanding, vision-language assessment, and fairness across domains and demographics.
  • Open, reproducible research. Code, datasets, and documentation shared openly to support governance-ready research.

For the full AIXpert vision, consortium, and funding details, see About.


System Architecture

Vector's responsible AI pipeline moves data through five stages: from raw inputs to governed, explainable outputs.

Synthetic Data Generation Fairness-aware multimodal data: images, VQA pairs, text scenes, and video, with demographic metadata and reproducible seeds.
↓
Multimodal Pipelines Parallel text, vision, video, and audio agents with attribution hooks and Risk-VQA for bias and toxicity detection.
↓
Agentic AI Evaluation Traceable planning and execution agents with RAG/memory, tool registry, and sandboxed task execution.
↓
Fairness Metrics + Explainability Statistical parity, equal opportunity, attribution and trace-based diagnostics, with disparity plots and explanation bundles.
↓
Responsible AI Insights Human-in-the-loop review, signed Governance Log (prompts, tool calls, safety decisions), and final explainable outputs.

Recent Updates

  • SONIC-O1 accepted at EMNLP 2026 (Main). Project page · Preprint · Code · Leaderboard.
  • DiagFlowBench accepted at EMNLP 2026 (Industry). Preprint.
  • AgentFinVQA accepted at NeurIPS 2026 AABA4ET. Project page · Preprint · Code.
  • Detecting Deception, Not Deepfakes accepted at NeurIPS 2026 TAE. Project page · Preprint.
  • Green AI Workshop (ECML-PKDD 2026). Shaina Raza delivered a keynote presenting DIA and the ICML 2026 Spotlight Position Paper. Program · Project.
  • Harness-aware evaluation. Survey treating each agent score as model, harness, environment, and evaluator. Project page · Preprint · Code.
  • Unified evaluation framework. Trustworthy evaluation across LLMs, agentic AI, and multimodal systems. Preprint.
  • Stress-testing efficient RAI evaluation accepted at NeurIPS 2026 TAE. Project page · Preprint · Code.
  • FairLens. Fairness benchmark for VLMs in high-stakes hiring, legal, and healthcare decisions. Project page · Preprint · Code · Dataset.
  • Deepfakes survey. Lifecycle survey of generation, distribution, and forensics in the foundation-model era. Project page · Preprint · Code.
  • UnBias-Plus. Bias detection and debiasing toolkit with paper, CLI, REST API, Python, and live demo. Project page · Code · Demo.
  • DIA (Data & Impact Accounting). Open-source toolkit that tracks the energy, water, and CO₂ footprint of open-source AI and its derivatives. Project page · Code · PyPI · Dashboard.
  • HumaniBench accepted at ACM TIST. Paper · Project page.
  • DIA in the press. Featured in Nosian Magazine's AI's Unasked Question: What Did That Cost?.
  • UnBias-Plus in the press. Independent coverage across CAN Health, ChannelLife, AI Loop, BornCity, TechTalent.ca, GlobeNewswire, and BNN Bloomberg.
  • MoE blog. Mixture of Experts: From Sparse Routing to Multimodal Deployment.
  • IASEAI 2026. Detecting and Reasoning About Bias in Multimodal Content.
  • AI4Good Lab 2026. Shaina Raza, PhD and Ahmed Y. Radwan presented UnBias-Plus and disinformation detection research at the AI4Good Lab 2026 Toronto cohort.
  • Toronto Machine Learning Summit. Ahmed Y. Radwan presented SONIC-O1 at the Toronto Machine Learning Summit (16-19 June 2026). Project page · Code · Leaderboard.
  • HAICON26 & Vector-Helmholtz Munich MOU. Shaina Raza, PhD presented at HAICON 2026 (8-11 June 2026, Munich) and Vector Institute signed an MOU with Helmholtz Munich's Computational Health Center.
  • AIXpert General Assembly: Barcelona 2026. The AIXPERT consortium met at the Barcelona Supercomputing Center (3-4 June 2026) to align the technical roadmap for year two.
  • The Peak Emerging Leaders 2026. Shaina Raza, PhD recognized in The Peak's Emerging Leaders 2026 in the Artificial Intelligence category.
  • AgentFinVQA. Auditable multi-agent pipeline for financial chart QA with traceable Model Evaluation Packets. Project page · Code.
  • FairSense-AgentiX. Agentic fairness and AI-risk analysis for text, images, and datasets. Project page · Code.

View full list


A snapshot of Vector's key contributions within AIXpert. Each project has its own repository, documentation, and quickstart.

  • UnBias-Plus

    AI-driven toolkit for bias detection and debiasing in text: biased spans, severity, reasoning, neutral replacements, and a full neutral rewrite for more trustworthy workflows.

    Project page · Code · PyPI

  • DIA (Data & Impact Accounting)

    Open-source toolkit to track energy, water, and CO₂ impact of open-source AI models and their derivatives, with a public dashboard.

    Project page · Code · PyPI · Dashboard

  • FairSense-AgentiX

    Agentic workflows for bias detection and risk assessment on text, images, and datasets: planning, tool use, self-critique, and telemetry-backed explanations.

    Project page · Code · PyPI

  • SONIC-O1

    Real-world benchmark for evaluating MLLMs on audio-video understanding, with a public leaderboard.

    Dataset · Code · Leaderboard

  • SONIC-O1 Multi-Agent

    Multi-agent framework for audio-video understanding with chain-of-thought reasoning, self-reflection, and temporal grounding.

    Code

  • Explainable Agentic Evaluation Framework

    Analyzes reasoning traces and interpretability of agentic AI across static and agentic settings.

    Code · Project page

  • Factual Preference Alignment (F-DPO)

    Factuality-aware preference learning to reduce LLM hallucinations without a separate reward model.

    Paper · Project page · Code

  • HumaniBench

    Fairness-focused vision-language benchmark evaluating foundation models across human-centric demographics.

    Project page · Paper

  • Agentic Transparency

    Survey and framework on interpretability, explainability, and governance of agentic AI systems.

    Project page

View all papers Projects & quickstarts


Citation

If you use any of our tools, datasets, or benchmarks, please cite the relevant work. BibTeX entries are available on each paper's entry in the Papers page.


Responsible AI Notice

This project may generate synthetic data containing demographic attributes for fairness research. These datasets are designed for controlled bias analysis and responsible AI evaluation only. They are not intended to represent or target real individuals. All data generation follows Vector Institute's responsible AI guidelines and AIXpert's ethical framework.


Have feedback or want to contribute? See the Team section on About and open an issue or pull request.