Data Marketplace

Use verified, licensed data with confidence. You can download right away or check the data through inquiry.

Sell your data on Flitto

Simply register your data and check its sale eligibility.

A total of 323 datasets
  • Pre-training DataText

    Low-Resource Language Parallel Translation Corpus Text Dataset

    A low-resource language parallel corpus dataset built by translating Korean written and spoken-style source texts into eight low-resource languages, including Vietnamese, Indonesian, Thai, and Hindi.

  • Frontier DataText

    Gulf Arabic Doctoral Exam Questions and Solutions

    A high-difficulty text dataset developed to train expert-level reasoning in LLMs based on doctoral examination questions and solutions in Gulf Arabic.

  • Frontier DataText

    Egyptian Arabic Doctoral Exam Questions and Solutions

    A high-difficulty text dataset developed to train expert-level reasoning in LLMs based on doctoral examination questions and solutions in Egyptian Arabic.

  • Frontier DataText

    Bengali Doctoral Exam Questions and Solutions

    A high-difficulty text dataset developed to train expert-level reasoning in LLMs based on doctoral examination questions and solutions in Bengali.

  • Frontier DataText

    Hindi Doctoral Exam Questions and Solutions

    A high-difficulty text dataset developed to train expert-level reasoning in LLMs based on doctoral examination questions and solutions in Hindi.

  • Frontier DataText

    Indonesian Doctoral Exam Questions and Solutions

    A high-difficulty text dataset developed to train expert-level reasoning in LLMs based on doctoral examination questions and solutions in Indonesian.

  • Frontier DataText

    Japanese Doctoral Exam Questions and Solutions

    A high-difficulty text dataset developed to train expert-level reasoning in LLMs based on doctoral examination questions and solutions in Japanese.

  • Frontier DataText

    Thai Doctoral Exam Questions and Solutions

    A high-difficulty text dataset developed to train expert-level reasoning in LLMs based on doctoral examination questions and solutions in Thai.

  • Pre-training DataAudio

    Voice Dataset with Embedded Foreign-Language Keywords

    A single-turn Korean speech dataset featuring sentences naturally embedded with foreign-language keywords, including English, Japanese, Chinese, Vietnamese, and Indonesian.

  • Pre-training DataAudio

    Malaysian English Multilingual Speech Dataset for Enterprise & Business Domains

    A single-turn Malaysian English speech dataset designed to improve multilingual interpretation and translation in enterprise and business domains.