LLM OrchestrationOpen Source AI & Machine Learning

Wildguard

Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Open Source

About

WildGuard is designed to classify the harmfulness of prompts and responses in user-model interactions. It can identify whether a prompt is harmful, assess the harmfulness of a response, and determine if a response is a refusal to answer. This tool is suitable for developers and researchers working with language models who need to ensure safety and compliance in conversational AI applications.

Open Source Health

Not enough history
Stars
138
Forks
15
License
Not stated
Last commit
2 years ago
Python

Related Categories