Zero-shot LLM toxicity classification on Civil Comments. Compares Integrated Gradients vs. attention attribution consistency, evaluates demographic fairness (SPD, EOpp), and provides an interactive Streamlit explorer with local and Gemini API backends.
nlp transformers attention zero-shot fairness bias-detection integrated-gradient huggingface streamlit toxicity-detection civil-comments llm interpretabiliity
-
Updated
Apr 30, 2026 - Python