JOIN THE GLOBAL VOICE AI GATHERING 👉
    Karya

    Karya

    Tech
    Dataset

    End-to-end AI data pipelines, evaluation benchmarks, and localized Indian datasets.

    Founded 2022Bangalore, INLinkedIn
    Karya banner

    About Karya

    Karya: Building AI for Real-World Complexity

    Karya designs and delivers end-to-end pipelines across data, evaluation, and deployment. Positioned as the AI Data and Evaluation Stack for India, the platform provides foundational datasets and evaluation benchmarks tailored to India's linguistic, cultural, and operational complexity across sectors like healthcare, agriculture, finance, law, education, and public services.

    Key Features

    • Data Collection: Provides custom data solutions for the frontier of AI, including domain-specific transcription, localized translation, and multimodal dataset creation at scale.
    • Conversational Speech Datasets: Offers large-scale conversational datasets across 22 official Indian languages.
    • Physical and Embodied AI: Features egocentric work and life datasets designed for physical-world and embodied AI systems.
    • Evaluation Benchmarks: Delivers national-scale evaluation frameworks, including Samiksha, the largest multilingual benchmark across 6 Indian languages for 17 models and 4 key domains.

    Use Cases

    • Domain-Specific AI Training: Utilizing localized translation and transcription for healthcare, agriculture, finance, law, education, and public services.
    • Multilingual Model Evaluation: Testing AI models using national-scale evaluation frameworks across multiple Indian languages.

    Getting Started

    Website: https://www.karya.in/

    Karya provides essential data collection and evaluation services to build better AI, focusing on the unique linguistic and operational needs of the Indian landscape.