On the missing data layer and a potential solution

작성자

카테고리:

← 피드로
arXiv cs.AI · Francis F Daniel, Mauro Iba~nez, Francis Perelman, Marian Basti · 2026-08-05 AI

[Submitted on 3 Aug 2026]

View PDF HTML (experimental)

Abstract:Latin America is missing two foundational layers of AI infrastructure: the dataset layer and the benchmark layer. This paper targets the dataset layer. The dataset layer faces two compounding problems: discovery and supply. Latin American AI datasets exist but are scattered across platforms with no shared index. Even with perfect indexing, the total volume would remain far below what frontier AI development requires. We propose DataHub: a task-first data infrastructure organized through the ontology /<task?>/<domain?>/<language?>, with mechanisms for dataset discovery, metadata, contribution, licensing, and reuse.

Submission history

From: Francis F Daniel [view email]
[v1] Mon, 3 Aug 2026 23:26:06 UTC (11 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.02949

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다