Coverage

Directory V0 covers crop disease and crop phenotyping image datasets for classification, object detection, and segmentation. A record counts as covered only when its minimum metadata, registered source, access state, collection geography state, and review information are present. Coverage does not mean OpenAgriData hosts the files.

Source and access

Curators create records manually or from reviewed CSV candidates. Every record links back to a registered source. Access is shown as available at source, request required, metadata only, restricted, or unavailable. License evidence and access mode are evaluated separately.

Classification and geography

Controlled terms describe crop or species, disease or trait, machine-learning task, modality, annotation type, acquisition method, format, license, and access. Country refers to where data was collected or observed, never the publisher address. Unknown and not provided remain explicit states.

Duplicate control

Before publication, curators compare stable identifiers, source-system IDs, canonical URLs, and normalized title plus publisher. SHA-256 is not used to decide dataset identity because the directory does not read or mirror source files. Potential matches remain explainable review candidates and are never merged automatically.

Updates

Registered source pages are checked using restrained HTTP requests and source metadata where available. Failures and changes create curator tasks; records are not silently deleted or overwritten.

AI assistance

Machine suggestions may help internal curation, but they do not become public facts without evidence and human acceptance. AI does not decide license rights, sensitive geography, or duplicate merges.