Create evaluation tasks, compare output quality across multiple models, and run RAG pipeline evaluations — all through the UI, no scripting required.