Financial QA Contrast-Set Generation and Evaluation Toolkit
Controlled perturbations and automatic validation for financial QA robustness testing
What it does#
Generates contrast-set instances from QA pairs derived from Volkswagen annual reports, applying controlled meaning-preserving and meaning-altering perturbations with vLLM and open-source Qwen, DeepSeek, and Llama models.
Validation pipeline#
An automatic validation pipeline combines NLI models and LLM-as-a-Judge evaluation to assess the semantic validity of generated contrast-set instances and filter out unreliable perturbations.
Human calibration#
A Streamlit-based annotation tool collects human annotations that calibrate the automatic evaluators, and the resulting contrast set is used to benchmark multiple QA models against each other.
last updated 2026.08.08
in-progress
Timeline
Sep 2025 – Present
Source
github.com/senemogluc/python-perturbation-generatorStack
PythonvLLMHugging FacePyTorch
Topics
NLPLLM EvaluationFinancial QA
On this page