• GlobalVQA
    Large-scale benchmark for evaluating global visual and textual perception in vision-language models across tasks involving scene recognition, embedded text, scene-text relationships, and global image understanding.
    [data]

  • LLM Access-Control Datasets
    Two large-scale permission-aware text-to-SQL datasets extending Spider and BIRD with role-based access-control policies, simulated users, database schemas, and ground-truth permit/deny decisions.
    [data]

  • Vision-to-Policy Access-Control Graph Dataset
    Multimodal dataset of access-control policy diagrams and structured ground-truth graphs for evaluating vision-language models on entity extraction, relation recovery, and policy-graph reconstruction.
    [data]

  • DoppelBot Human-LLM Collaboration Dataset
    Anonymized interaction logs and experimental data from live collaborative tasks in which middle-school students interacted with human and LLM-controlled agents and attempted to identify the AI participants.
    [data]

  • Bike Frames
    Dataset of 31,480 cycling-related news headlines, including 1,500 human-annotated examples capturing cyclist perception, accident reporting, and attribution of fault.
    [data]

  • IoT-SQL
    Cybersecurity dataset combining an IoT relational database, natural-language-to-SQL queries, and labeled network traffic for research on text-to-SQL, IoT threat detection, and multimodal security analysis.
    [data]

  • BioASR-NER
    Biomedical speech and named entity recognition dataset containing nearly 2,000 clean and noisy recordings and transcripts for studying the performance gap between automatic speech recognition and downstream biomedical NLP systems.
    [data]

  • Chemical NER Gender Bias Dataset
    Benchmark for evaluating demographic performance disparities in chemical named entity recognition, including an annotated Reddit corpus with self-identified gender information and large-scale synthetic evaluation data.
    [data]

  • GeoOLID
    Geographically diverse offensive-language dataset containing more than 14,000 examples collected across 15 U.S. cities, with labels for geographic location and offensive language.
    [data]

  • WallStreetBets Intent and Support Dataset
    Annotated Reddit dataset centered on the GameStop trading phenomenon for studying users' intent to purchase stocks and their support for coordinated community actions.
    [data]