
LLM Response Evaluation Toolkit
A rubric-driven workflow for scoring large language model responses on accuracy, reasoning quality, and guideline adherence. Includes prompt templates, a scoring schema, and reporting that surfaces where a model drifts.
- Python
- Prompt Engineering
- LLM Evaluation

