Sustainable Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions

Ahmed El Kady1 Aravind Narayanan1 Rehana Riaz2 Yani Ioannou3 Shaina Raza1,*

1Vector Institute 2Independent researcher 3University of Calgary

* Corresponding author: shaina.raza@vectorinstitute.ai
Paper · Coming soon arXiv · To be added Code

Can responsible-AI evaluation use substantially less compute without changing the conclusions we draw from it?

We measure evaluation quality and operational footprint together, rather than treating efficiency as a proxy for trustworthy results.

How we evaluate

Swipe to follow the workflow

Horizontal workflow from frozen BBQ and BBQ-V benchmarks through evaluation conditions, model inference, and answer and reasoning. Evaluation quality and operational footprint are measured together and compared with the M0 full BF16 baseline.

Each intervention is compared with the same full-benchmark BF16 baseline (M0).

What we evaluate

Benchmarks

BBQ and BBQ-V

Frozen evaluation membership from the original benchmarks.

Models

  • Qwen2.5-VL-7B
  • Qwen3-VL-30B-A3B
  • Gemma-4-12B-it

Efficiency interventions

  • M0 BF16 baseline
  • M1 larger batching
  • M2 / M3 INT8 / INT4
  • M4 / M5 reduced benchmark / combinations

Measured outcomes

Quality
Accuracy, Bias Score / Bias Present, Reasoning Quality

Footprint
Inference time, measured GPU energy, estimated CO2e, estimated water use

Compute savings change what the benchmark tells us

01

Batching is the reliable first lever

Across the evaluated models and benchmarks, it largely preserved quality and usually lowered energy use.

02

INT8 is not an energy shortcut here

It preserved average quality, but raised runtime and energy use in every evaluated setting.

03

Smaller benchmarks save most, with a limit

They gave the most consistent savings, but conclusions became less stable at the smallest sizes.

Energy, accuracy, and bias must be read together. Highlighted M0 and M4 points show the full baseline and reduced-benchmark conditions.

Swipe to inspect the figure

Measured GPU energy versus accuracy for BBQ and BBQ-V. Marker shape denotes Qwen2.5, Qwen3-MoE, or Gemma 4; highlighted M0 and M4 points show the full baseline and reduced-benchmark conditions.

The horizontal axis is logarithmic. Darker markers indicate lower Bias Present values.

Citation

BibTeX: To be added

% To be added

Acknowledgments

Resources used in preparing this research were provided, in part, by the Province of Ontario, the Government of Canada through CIFAR, and companies sponsoring the Vector Institute.

This research was funded by the European Union's Horizon Europe research and innovation programme under the AIXPERT project (Grant Agreement No. 101214389).

Questions and collaborations

For questions or collaborations, please open an issue in this repository or contact the corresponding author at shaina.raza@vectorinstitute.ai, as listed in the paper.