Collection
AI Safety & Alignment
The AI Safety & Alignment collection tracks 14 curated open-source projects, 11 of them live on TrendingRepo right now — ranked by a cross-source momentum score blending GitHub star velocity with mentions on Hacker News, X, Bluesky, Product Hunt and Dev.to.
Live · top 14 repos · sorted by momentum across 24H
LIVE · 11m| # | Repository | Stars | 24h | 7d | 30d | Trend | Mentions | Actions |
|---|---|---|---|---|---|---|---|---|
| 01 | NVIDIA/garak the LLM vulnerability scanner | 8.6K | +9+0.1% | +72+0.8% | +356+4.3% | |||
| 02 | NVIDIA-NeMo/Guardrails NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems. | 6.8K | +4+0.1% | +49+0.7% | +240+3.7% | |||
| 03 | meta-pytorch/captum Model interpretability and understanding for PyTorch | 5.7K | +1+0.0% | +4+0.1% | +21+0.4% | |||
| 04 | shap/shap A game theoretic approach to explain the output of any machine learning model. | 25.6K | +4+0.0% | +25+0.1% | +104+0.4% | |||
| 05 | fairlearn/fairlearn A Python package to assess and improve fairness of machine learning models. | 2.3K | +1+0.0% | +4+0.2% | +11+0.5% | |||
| 06 | guardrails-ai/guardrails Adding guardrails to large language models. | 7.2K | +1+0.0% | +22+0.3% | +150+2.1% | |||
| 07 | protectai/llm-guard The Security Toolkit for LLM Interactions | 3.2K | +1+0.0% | +8+0.3% | +79+2.5% | |||
| 08 | openai/evals Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks. | 19K | +1+0.0% | +64+0.3% | +250+1.3% | |||
| 09 | Trusted-AI/AIF360 A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in datasets and models. | 2.8K | — | +2+0.1% | +16+0.6% | |||
| 10 | llm-attacks/llm-attacks Universal and Transferable Attacks on Aligned Language Models | 4.7K | +2+0.0% | +7+0.1% | +39+0.8% | |||
| 11 | microsoft/responsible-ai-toolbox Responsible AI Toolbox is a suite of tools providing model and data exploration and assessment user interfaces and libraries that enable a better understanding of AI systems. These interfaces and libraries empower developers and stakeholders of AI systems to develop and monitor AI more responsibly, and take better data-driven actions. | 1.8K | +1+0.1% | +7+0.4% | +23+1.3% | |||
| 12 | microsoft/presidio No description published. | 0 | — | — | — | |||
| 13 | openai/safety-gym No description published. | 0 | — | — | — | |||
| 14 | Trusted-AI/AIX360 No description published. | 0 | — | — | — |