File-system storage adapter

The file-system adapter is a concrete implementation of the storage protocols described in the Data storage page, section StorageAdapter interface.

The FsStorageAdapter class implements the StorageAdapter protocol, relying on FsStorageReader / FsStorageWriter for reading/writing and FsProviderReader / FsProviderWriter for provider-level operations.

This adapter has options that are detailed below.

The physical data model

Although using human-readable file formats like JSON, JSON Lines and TSV, it can be considered as a black box, meaning that reading and writing data should be done through the storage adapter.

The data format of the files is designed to optimize reads and writes and can differ from the Data model classes.

To represent the content of those files in-memory, another set of model classes is defined in dbnomics_toolbox.storage.adapters.filesystem.model. We can talk about those classes as the physical model, whereas the classes of dbnomics_toolbox.model can be called the domain model.

When calling load and save methods on FsProviderReader or FsProviderWriter, they are responsible for both:

  • transforming objects of the domain model from and to objects of the physical model,

  • knowing the path of the files that are read and written.

Provider metadata

The file-system adapter stores provider metadata in a file named provider.json directly at the root of the provider directory.

The ProviderMetadata domain model class is mapped to the ProviderJson physical model class.

See also: ProviderReader.read_provider_metadata and ProviderWriter.write_provider_metadata.

Category tree

The file-system adapter stores the category tree in a file named category_tree.json directly at the root of the provider directory.

The CategoryTree domain model class is mapped to the CategoryTreeJson physical model class.

See also: ProviderReader.read_category_tree and ProviderWriter.write_category_tree.

Datasets and series

The file-system adapter stores each dataset in a dedicated directory named after the dataset code. For example, the dataset INSEE/IPC-2015 is stored in the directory IPC-2015 at the root level of the provider directory.

The way datasets and series are stored depends on the chosen dataset layout, the different variants being described in the next section.

Dataset layout

The storage URI of the file-system adapter accepts a variant parameter that takes one of those values: jsonl or tsv (e.g. filesystem:insee-converted-data?variant=jsonl).

When instantiating a FsStorageAdapter, the variant will be detected by looking at the already written files, if any. If detection could not be done, for example because no file has been written yet, then the jsonl variant will be used by default for future writes. This allows clients to open any directory with the file-system adapter without knowing by advance which variant is used.

TSV variant

The TSV variant mainly uses tab-separated values files to store series.

When using the TSV variant, the following files are created for each dataset, in a sub-directory named after the dataset code.

The file {dataset_code}/dataset.json contains dataset metadata coming from the DatasetMetadata model class, and series metadata coming from the Series model class under the series property. The TsvDatasetJson physical model class represents the contents of this file.

For each series, observations and their attributes are stored in a TSV file named after the series code. For example, the series INSEE/IPC-2015/A.IPC.SO.00.00.INDICE.ENSEMBLE.FE.SO.BRUT.2015.FALSE is stored in the file IPC-2015/A.IPC.SO.00.00.INDICE.ENSEMBLE.FE.SO.BRUT.2015.FALSE.tsv.

A simple TSV file looks like this:

PERIOD VALUE
2000 19
2001 NA
2002 22

Observation attributes can be stored in additional columns:

PERIOD VALUE OBS_STATUS
2000 19
2001 NA
2002 22 E

Historically the TSV variant was the only existing variant, but since it uses one TSV file per series, this could lead to a huge number of files for some datasets. Given that in file-systems, any file takes a minimum of one block (e.g. 4kb), this is not optimal for datasets having a huge number of small series. Quickly the disk was full due to those millions of small files.

Also the fact that the file is named after the series code can lead to file names that are too long for the runtime environment (i.e. file-system, operating system).

JSON Lines variant

The JSON Lines variant mainly uses JSON Lines files to store series.

This variant was introduced to circumvent the issues and limitations of the TSV variant.

The file {dataset_code}/dataset.json contains dataset metadata coming from the DatasetMetadata model class. The JsonLinesDatasetJson physical model class represents the contents of this file.

The file {dataset_code}/series.jsonl contains all the series of the dataset, including metadata and observations coming from the Series model class. The JsonLinesSeriesItem physical model class represents each line of this file.

Example of series.jsonl (only a minimal sample is shown here):

{"code":"M.LB.B.TTP.SA","dimensions":{"frequency":"M","seasonally_adjusted":"SA","sex":"B","subject":"LB","unit":"TTP"},"observations":[["PERIOD","VALUE"],["1953-01",4122],["1953-02",4001],["1953-03",4008]]}
{"code":"M.NILF.F.TTP.NSA","dimensions":{"frequency":"M","seasonally_adjusted":"NSA","sex":"F","subject":"NILF","unit":"TTP"},"observations":[["PERIOD","VALUE"],["1972-01","NA"],["1972-02","NA"],["1972-03","NA"],["1972-04","NA"]]}

Note that in the above example the series are already sorted alphabetically by code (M.LB.B.TTP.SA < M.NILF.F.TTP.NSA). This is a constraint of the JSON Lines format: series must appear in series.jsonl sorted alphabetically by their code.

Sorting series enables efficient lookups via JsonLinesSeriesOffsetIndex — the file-system adapter can jump directly to the byte offset of a given series without scanning the entire file.

However, sorting a potentially huge number of series entirely in memory is impractical. To avoid loading everything into memory, the writer uses a temporary SQLite database (series.sqlite) as an intermediate store: each series is serialized to JSON and indexed by its code in SQLite, then retrieved in sorted order via ORDER BY. This is implemented by BlobIndexer.

Provider Writer Sessions

FsProviderWriter supports ProviderWriterSession for atomic writes. See Data storage, section Provider Writer Sessions, for the general concept.

When using the file-system adapter, a session writes all changes to a temporary .sessions/{session_id}/ directory. Only when ProviderWriterSession.commit is called, the files are moved to the target storage directory atomically.

Single provider or multiple providers

By design, the storage protocols are able to load and save data belonging to several providers:

provider_reader = storage_reader.get_provider_reader(parse_provider_code("INSEE"))
provider_reader2 = storage_reader.get_provider_reader(parse_provider_code("IMF"))

However, by design also, fetchers are dedicated to a single provider.

For convenience during fetcher development, a storage directory can be configured to handle a single provider (rather than requiring a top-level directory per provider). This avoids the extra directory nesting when working on one fetcher at a time.

The file-system adapter reads and writes data from/to a directory, which can be used in 2 different modes, according to the single_provider boolean parameter of the storage URI.

With single_provider=false (the default), the adapter works in multiple provider mode. The path of the storage URI will be used as a base directory containing one top-level directory per provider.

With single_provider=true, the adapter works in single-provider mode. The path of the storage URI will be used as a directory containing data of a single provider only (the same as the top-level directories of the multiple provider mode).

Even when being used in the single-provider mode, the provider codes must be given to the methods of the storage protocols in order to respect the common interface.

Example for multiple provider mode:

from dbnomics_toolbox.model.provider_metadata import ProviderMetadata
from dbnomics_toolbox.model.identifiers import parse_provider_code
from dbnomics_toolbox.storage.adapters.factories import create_storage_adapter_from_uri
from dbnomics_toolbox.storage.uris import Uri

multiple_provider_storage = create_storage_adapter_from_uri(Uri.parse("filesystem:///multiple_provider_data"))
multiple_provider_writer = multiple_provider_storage.get_writer().get_provider_writer(parse_provider_code("EUROSTAT"))
multiple_provider_writer.write_provider_metadata(
    ProviderMetadata.create(
        code="EUROSTAT",
        name="Eurostat",
        # Cf https://en.wikipedia.org/wiki/ISO_3166-1_alpha-2
        region="EU",
        terms_of_use="https://ec.europa.eu/eurostat/web/main/help/copyright-notice",
        website="https://ec.europa.eu/eurostat",
    )
)
multiple_provider_storage.get_writer().get_provider_writer(parse_provider_code("INSEE")).write_provider_metadata(
    ProviderMetadata.create(
        code="INSEE",
        name="Institut national de la statistique et des études économiques",
        # Cf https://en.wikipedia.org/wiki/ISO_3166-1_alpha-2
        region="FR",
        terms_of_use="https://www.insee.fr/fr/information/2381863",
        website="https://www.insee.fr/",
    )
)
$ tree multiple_provider_data
multiple_provider_data
├── eurostat-json-data
│   └── provider.json
└── insee-json-data
    └── provider.json

Example for single-provider mode:

from dbnomics_toolbox.model.provider_metadata import ProviderMetadata
from dbnomics_toolbox.model.identifiers import parse_provider_code
from dbnomics_toolbox.storage.adapters.factories import create_storage_adapter_from_uri
from dbnomics_toolbox.storage.uris import Uri

single_provider_storage = create_storage_adapter_from_uri(
    Uri.parse("filesystem:///single_provider_data?single_provider=true")
)
single_provider_storage.get_writer().get_provider_writer(parse_provider_code("INSEE")).write_provider_metadata(
    ProviderMetadata.create(
        code="INSEE",
        name="Institut national de la statistique et des études économiques",
        # Cf https://en.wikipedia.org/wiki/ISO_3166-1_alpha-2
        region="FR",
        terms_of_use="https://www.insee.fr/fr/information/2381863",
        website="https://www.insee.fr/",
    )
)
$ tree single_provider_data
single_provider_data
└── provider.json

In a fetcher, the convert CLI expects a directory or a storage URI as its second command-line argument. If it’s a directory, it is turned into a storage URI like filesystem:{directory}?single_provider=true.