promptfoo

    promptfoo/promptfoo

    #40 this week

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

    devops
    llm
    testing
    ci
    ci-cd
    cicd
    evaluation
    evaluation-framework
    TypeScript
    MIT
    24.0K stars
    2.2K forks
    24.0K GitHub watchers
    Updated 8/7/2026
    View on GitHub

    Backblaze Generative Media Hackathon

    Build the next generation of AI media apps with Genblaze, stored on Backblaze B2. $10,000 in prizes.

    Enter the hackathon

    Loading star history...

    Use Cases & Benefits

    • Provides a developer-friendly tool to test, evaluate, and secure prompts, agents, and retrieval-augmented generation (RAG) applications for large language models.
    • Enables local, private, and automated testing with flexible support for multiple LLM APIs and seamless CI/CD integration to ensure reliable and secure AI deployments.
    • Use for automated evaluation and comparison of different LLMs like GPT, Claude, and Gemini to select the best model for your application.
    • Use for red teaming and vulnerability scanning to identify and mitigate security risks in AI prompt engineering and LLM-based applications.
    • Use for integrating prompt and model testing into CI/CD pipelines to automate quality checks and code scanning for LLM-related security and compliance issues.

    About promptfoo

    Promptfoo: LLM evals & red teaming

    npm npm GitHub Workflow Status MIT license Discord

    promptfoo is a developer-friendly local tool for testing LLM applications. Stop the trial-and-error approach - start shipping secure, reliable AI apps.

    Website · Getting Started · Red Teaming · Documentation · Discord

    Quick Start

    # Install and initialize project
    npx promptfoo@latest init
    
    # Run your first evaluation
    npx promptfoo eval
    

    See Getting Started (evals) or Red Teaming (vulnerability scanning) for more.

    What can you do with Promptfoo?

    • Test your prompts and models with automated evaluations
    • Secure your LLM apps with red teaming and vulnerability scanning
    • Compare models side-by-side (OpenAI, Anthropic, Azure, Bedrock, Ollama, and more)
    • Automate checks in CI/CD
    • Review pull requests for LLM-related security and compliance issues with code scanning
    • Share results with your team

    Here's what it looks like in action:

    prompt evaluation matrix - web viewer

    It works on the command line too:

    prompt evaluation matrix - command line

    It also can generate security vulnerability reports:

    gen ai red team

    Why Promptfoo?

    • 🚀 Developer-first: Fast, with features like live reload and caching
    • 🔒 Private: LLM evals run 100% locally - your prompts never leave your machine
    • 🔧 Flexible: Works with any LLM API or programming language
    • 💪 Battle-tested: Powers LLM apps serving 10M+ users in production
    • 📊 Data-driven: Make decisions based on metrics, not gut feel
    • 🤝 Open source: MIT licensed, with an active community

    Learn More

    Contributing

    We welcome contributions! Check out our contributing guide to get started.

    Join our Discord community for help and discussion.

    Discover Repositories

    Search across tracked repositories by name or description