š«±š¾āš«²š¼ Human Iteration on AI evals: Creating "Tasks" for Better Collaboration & Quality Control
Struggling to manage and assign work for AI evals? š¤ We're introducing the "Tasks" feature to help you and your team effectively manage and distribute work related to LLM evaluation and security. What you will learn: - Task Dashboard: Get a centralised overview of all tasks, including priority, status, and assignments. Filter by your tasks or unassigned items. - Review Workflow: Create and assign tasks directly within Vulnerability Scan results and Evaluation Runs to review specific items or test cases. - Quality Gates: Use the Draft conversation status to prevent a conversation from being reused in subsequent evaluations until all related tasks are resolved and published. - Higher Quality Datasets: Tasks ensure no evaluations are missed, leading to higher quality and more consistent evaluation datasets. Start managing, distributing, and controlling your LLM evaluation configurations with your team! Useful resources: - Free trial: https://www.giskard.ai/contact Prevent AI failures, don't react to them.