General Preference Evaluation Chat Arena Multi-Turn Dataset
A Korean-based Chat Arena preference evaluation multi turn dataset built by collecting human preferences for AI model responses.
Use verified, licensed data with confidence. You can download right away or check the data through inquiry.
A Korean-based Chat Arena preference evaluation multi turn dataset built by collecting human preferences for AI model responses.
A Korean-based long context benchmark text dataset built to evaluate AI models' ability to process and reason over extended contexts.
A high-difficulty Korean reasoning benchmark dataset built according to global HLE standards, including complex expert-level reasoning problems.
A high-difficulty Korean HLE reasoning and benchmark dataset built with Korea-specific expert-level reasoning problems.
A benchmark text dataset developed for Arena evaluation based on Gulf Arabic.
A benchmark text dataset developed for Arena evaluation based on Egyptian Arabic.
A benchmark text dataset developed for Arena evaluation based on Bengali.
A benchmark text dataset developed for Arena evaluation based on Hindi.
A benchmark text dataset developed for Arena evaluation based on Indonesian.
A benchmark text dataset developed for Arena evaluation based on Japanese.