Introduction
What is Dump Things Service?
Dump Things Service is a central component of a system that allows cooperative authoring of linked data (linked data refers to data that is structured according to a schema which is defined in LinkML). It supports a workflow where authors submit records into inboxes and curators move records from inboxes to a curated area. The curated area holds the “official”, high-quality, content.
A curation workflow moves proposed changes from the inboxes into the curated area. Curators inspect proposed changes, accept, modify, or reject them. Curation can be performed manually or by automated agents. The service provides tools and APIs to support both (all functionality is exposed via an HTTP API).
The main elements
A Dump Things Service instance (usually a process that runs on a server) supports multiple Collections. Each collection has a name and an associated schema. All records that are stored in the collection must conform to a schema that is associated with the collection. Every records has at least a persistent identifier (PID) that is a globally unique identifier (more about schemas below).
Clients that work with records on a Dump Things Service usually have one or more of the following roles:
Readers: users that read data from the curated area of a collection.
Authors: users that can submit new records for inclusion into the curated area of a collection. Their submissions are stored in author-specific inboxes. Authors can usually also read data from the collection.
Curators: users that can curate the collection. Curators have read and write access to the curated area of a collection. They also have read and write access to all inboxes. Curators usually move records from inboxes to the curated area, probably with editorial changes.
Besides an API that allows access to records, there is another API that allows access to the configuration of a Dump Things Service and permits dynamic configuration of an Dump Things Service instance at runtime. Entities with an Admin role can use this API to change the configuration of a Dump Things Service at runtime.
Interaction with Dump Things Service
Depending on the role of a user, there are different interaction patterns with a Dump Things Service. Those are: Reading, Writing, and Curating.
Reading
The most elementary interaction is to read records from a collection. Depending on the role of a user, reading will either return only records from the curated area or records from the curated area and from the user’s inbox.
A user with read-access for a collection will be able to retrieve any record that is stored in the curated area of a collection.
A user with read- and write-access for a collection, i.e., an author, will be able to retrieve any record that is stored in the curated area of a collection or in the inbox of the user. If records from the curated area and the inbox have identical PIDs, the record from the inbox will be returned.
Writing
A user with write-access for a collection, i.e., an author, can submit new records to the collection. If the records confirm the schema of the collection, they will be stored in the inbox of the user. Note: write-access does not allow to move records from the inbox to the curated area.
Curating
The purpose of curating is to moves proposed changes from the inboxes of a collection to the curated area of the collection.
Curator-rights allow a user to read, write, and delete records from the curated area directly. The user can also read, write, and delete records from any inbox of the collection.
Curators use the API to inspect proposed changes, accept, modify, or reject them. The client library dump-things-pyclient (see The dump-things ecosystem) provides tools that support curation processes.
Administration
Administrators do not operate on records. Administrators can inspect and modify the configuration of a Dump Things Service instance at runtime. An administrator can modify three configuration elements:
Collections: add, remove, or modify collections.
Tokens: add, remove, or modify tokens.
Administrator token: add. remove, or modify administrator tokens.
Records and Schemas
Records that are stored in a collection are instances of the classes that are
defined in the schema of the collection. The schema can be freely defined,
but all record-classes, i.e., classes whose instances are can be stored in a
collection, must be a subclass of the class Thing, which is
defined in the schema things-schema. The crucial property of the class
Thing is the persistent identifier (PID). The PID is an IRI
(Internationalized Resource Identifier), a globally unique
identifier that identifies the record.
The dump-things ecosystem
Dump Things Service is part of a larger ecosystem of tools and services that support the shared authoring and curation of linked data. Currently there are two other components that are built to work with it:
shacl-vue: a web application that allows users to view and edit linked data records stored in a Dump Things Service instance.
dump-things-pyclient: a Python client library that simplifies authoring of Dump Things clients in Python. It also provides a CLI-command to interact with a Dump Things Service from a shell.
Dump Things Structure
Workflow
Dump Things Service was built to support cooperative authoring of linked data. It supports a workflow where authors submit records into inboxes and curators move records from inboxes to a curated area.
Dump Things Service supports distributed data acquisition while still controlling the quality of collection data. This is done by distinguishing between two areas for records:
Curated area: represents the “official”, high-quality, collection content.
Incoming area (user-specific inboxes): where users propose new records or changes to existing records.
A curation workflow moves proposed changes from the incoming area into the curated area. Curators — entities with specific privileges — can inspect proposed changes, accept, modify, or reject them. Curation can be performed by persons or by automated agents. The service provides tools and APIs to support both options.
Collections
Dump Things Service supports on the highest level multiple collections. Each collection has a name and an associated schema. Collections have inboxes for all authors and a curated area.
User groups
We distinguish he following user groups:
Readers: user that read data from a collection.
Authors: users that can submit new records to a collection. Authors can usually also read data from the collection.
Curators: users that can curate the collection. Curators have read and write access to the curated area of a collection. They also have read and write access to all inboxes. Curators usually move records from inboxes to the curated area, probably with editorial changes.
Dump Things Service is an implementation of a service that allows storing and retrieving linked data. It builds the backend for linked-data applications. Linked-data applications can retrieve data as JSON records or as TTL (Turtle) documents from a Dump Things Service instance. Multiple Dump Things Service instances currently serve as backends for shacl-vue.
All data items stored on a Dump Things Service must be structured according to a defined schema. Schemata are defined in LinkML. The service supports multiple collections. Each data item is stored in a collection. Each collection has a name and an associated schema. The service dynamically reads the respective schemata and creates validation code for data items. This ensures that only items that adhere to the schema can be stored in the respective collection.
Curated Area and Incoming Area
The Dump Things Service supports distributed data acquisition while still controlling the quality of collection data. This is done by distinguishing between two areas:
Curated area: represents the “official”, high-quality, collection content.
Incoming area (user-specific inboxes): where users propose new records or changes to existing records.
A curation workflow moves proposed changes from the incoming area into the curated area. Curators — entities with specific privileges — can inspect proposed changes, accept, modify, or reject them. Curation can be performed by persons or by automated agents. The service provides tools and APIs to support both options.
Tokens and Permissions
To submit a record to a collection, a token is required. Tokens carry:
Read and write permissions for the incoming areas of collections.
Read permissions for the curated area of a collection.
A submitter ID that is annotated to each submitted record.
A token-specific zone in the incoming area.
Multiple tokens can share the same zone, allowing multiple submitters to work together on the same incoming area.
Key Features
Multiple storage backends:
record_dir,record_dir+stl,sqlite, andsqlite+stl.Flexible authentication: config-file-based tokens and Forgejo-based authentication.
Audit logs: track time, content, author, and curator for every change.
LinkML-based validation: ensure data conforms to a defined schema.