Wildguard
Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
About
WildGuard is designed to classify the harmfulness of prompts and responses in user-model interactions. It can identify whether a prompt is harmful, assess the harmfulness of a response, and determine if a response is a refusal to answer. This tool is suitable for developers and researchers working with language models who need to ensure safety and compliance in conversational AI applications.
Open Source Health
- Stars
- 138
- Forks
- 15
- License
- Not stated
- Last commit
- 2 years ago
Related Categories
Vendor
Allenai
Publisher of Dont Stop Pretraining, Objaverse Xl and Procthor
Quick Links
Open Source
More by Allenai
Related Products
Dont Stop Pretraining
Code associated with the Don't Stop Pretraining ACL 2020 paper
More from this vendor
The Algorithm
Source code for the X Recommendation Algorithm
Shared categories
Recommenders
TensorFlow Recommenders is a library for building recommender system models using TensorFlow.
Shared categories
RPG_KDD2025
This repository provides the code for implementing RPG described in our KDD'25 paper "Generating Long Semantic IDs in Parallel for Recommendation".
Shared categories
Pairec
A Go web framework for quickly building recommendation online services based on JSON configuration.
Shared categories