Promptfoo vs sizeless: Features, Pricing & Which Is Better (2026)
A side-by-side comparison of Promptfoo and sizeless — features, pricing, and ideal use cases — to help you decide which AI tool fits your workflow.
Promptfoo
Promptfoo
CLI and web tool for testing, evaluating, red‑teaming, and monitoring LLM prompts and outputs to catch regressions and vulnerabilities.
Key features
- Red-Teaming & Vulnerability Scanning: Declarative red‑team tests and automated scans to surface prompt injections, unsafe completions, and other model security risks across providers.
- Evaluations & Regression Detection: Run reproducible eval suites and compare outputs before/after changes to detect regressions, with CI/CD and GitHub Action integration for automated checks on PRs.
- Multi-Provider Model Comparison: Execute the same tests across multiple model providers and families (e.g., OpenAI, Claude, Gemini, Llama) to compare quality and safety consistently.
- CLI and Web UI: Command‑line tools for running tests and a web 'view' UI to inspect prompts, final rendered prompts, outputs, and structured results in tabular form.
- Declarative Configs & Templating: Use promptfooconfig.yaml with Nunjucks templating and custom filter plugins to generate complex prompts and test permutations programmatically.
- Extensible Provider & Plugin System: Add or customize providers, local execution, or custom filters (JS/Python) to adapt tests to specific stacks or private model endpoints.
- Docker Distribution & Local Execution: Official container images and local execution modes enable isolated, reproducible runs and CI friendliness.
- GitHub & CI Integrations: Official GitHub Action and CI-friendly tooling to automatically post evaluations on PRs and enforce prompt quality gates.
- Command‑line interface and library for running declarative evals and tests
- Red‑teaming and vulnerability scanning for LLM outputs
- Declarative configuration via promptfooconfig.yaml (prompts, providers, filters, tests)
- Support for templated prompts using Nunjucks and custom filter modules
- Providers for multiple model backends (OpenAI and others; compare GPT, Claude, Gemini, Llama, etc.)
- Docker images published to GHCR (multi‑arch support: linux/amd64, linux/arm64, etc.)
- Web UI (src/app) that integrates with `promptfoo view` for inspecting outputs and final prompts
- CI/CD integrations including an official GitHub Action for evals on PRs
- Developer productivity features: live reload, caching, npm scripts for local dev
- Configurable Python executable (PROMPTFOO_PYTHON) and language‑agnostic test data (supports Python, JavaScript, others)
Best for
- Red‑teaming LLM integrations to find prompt injections, unsafe outputs, and info‑leakage before release.
- Regression testing in CI to automatically detect when a prompt or model update degrades output quality or safety on pull requests.
- Comparing model performance across providers and model families to choose the best model for a given task or guardrail requirements.
- Building test-driven prompt development workflows where prompts are versioned, evaluated, and iterated using reproducible eval suites.
- Adding automated before/after eval diffs on GitHub PRs to give reviewers quantitative and qualitative signal about prompt edits.
- Validating agents, RAG pipelines, and LLM apps end‑to‑end by running scenario-based tests and inspecting final rendered prompts and outputs.
- Test‑driven prompt engineering and automated evaluation of model outputs
- Red‑teaming and security testing of language model behavior
- Regression testing of prompts and model changes via CI/CD and GitHub Actions
- Comparing performance across multiple model providers
- RAG (retrieval augmented generation) and agent testing in local/dev environments
- Integrating automated evals into PR workflows to produce before/after views of prompt edits
sizeless
sizeless
Turns a smartphone video of an open trench into a centimetre-accurate 3D point cloud, CAD as-built plan and GIS-ready digital twin of buried utilities.
Key features
- Smartphone capture: Field crews record an open trench with a standard iPhone Pro — no specialist scanning hardware and no separate surveying appointment
- Centimetre-accurate point clouds: Reconstruction algorithms developed at ETH Zurich build a high-resolution 3D point cloud of the excavation from the video alone
- Standards-compliant CAD output: Generates as-built plans in DWG and DXF, with couplings and pipe runs identified and measurements simplified
- 3D digital twin and GIS export: Produces a model of the pipe route including building entries that drops into existing GIS systems
- Works without GPS: Captures basement sections and building entry points where GNSS-based surveying fails
- Immediate backfilling: Because capture takes minutes, trenches close right after filming instead of waiting on a survey crew
- Documentation in about 72 hours: Complete records arrive weeks earlier than conventional surveying, enabling prompt connection billing
- Third-party utility capture: Records crossing utilities and as-laid geometry as unbroken 3D evidence, replacing hand sketches
Best for
- A utility network operator documenting residential service connections without booking a surveyor for every site
- A contractor closing a trench the same day instead of leaving it open pending a survey appointment
- Capturing a building entry point in a basement where GPS-based surveying cannot get a fix
- A district heating project producing as-built DWG plans for regulatory sign-off
- Spotting a laying error in the 3D point cloud before backfilling, while the fix is still cheap
- Feeding as-built pipe geometry into a GIS system for long-term network maintenance planning
