cxgin.dev

Financial QA Contrast-Set Generation and Evaluation Toolkit

Controlled perturbations and automatic validation for financial QA robustness testing

What it does#

Generates contrast-set instances from QA pairs derived from Volkswagen annual reports, applying controlled meaning-preserving and meaning-altering perturbations with vLLM and open-source Qwen, DeepSeek, and Llama models.

Validation pipeline#

An automatic validation pipeline combines NLI models and LLM-as-a-Judge evaluation to assess the semantic validity of generated contrast-set instances and filter out unreliable perturbations.

Human calibration#

A Streamlit-based annotation tool collects human annotations that calibrate the automatic evaluators, and the resulting contrast set is used to benchmark multiple QA models against each other.

last updated 2026.08.08
in-progress
Timeline
Sep 2025 – Present
Source
github.com/senemogluc/python-perturbation-generator
Stack
PythonvLLMHugging FacePyTorch
Topics
NLPLLM EvaluationFinancial QA
On this page