Skip to main content
An extraction job is an asynchronous request to read data out of one or more registered databases. The Manager accepts the job and queues it. A Worker runs the extraction and stores the result. This page explains the request you send and the path the job takes.

The request


A job carries two parts: dataRequest and metadata.
mappedFields is the heart of the request. It maps each datasource config name to the tables you want, and each table to the fields you want. metadata.source is required. It names the product that owns the job. The Manager rejects a request without it. The Manager also rejects the request when a referenced connection belongs to a different product, or belongs to no product at all. Internal datasources carry no product, so this check skips them.

Field projection


Name the fields you want, or ask for all of them:
  • A list of field names extracts exactly those fields.
  • The single entry ["*"] extracts every column of the table. The wildcard must be the only entry in the list.
Fetcher validates every named field against the discovered schema before it queries. An unknown field name fails the job and names the offending fields in the error. A table with no valid field also fails the job.

Multiple datasources, tables, and schemas


One job can read from several datasources at the same time. Each datasource can contribute several tables. Each table can sit in a different schema. Write a table name in one of two forms:
  • Unqualifiedaccounts. Fetcher reads it from the engine default schema: public on PostgreSQL, dbo on SQL Server.
  • Schema-qualifiedaccounting.invoices. The prefix before the dot is the schema. On Oracle it is the owner.
Fetcher collects the distinct schema prefixes in mappedFields and discovers only those namespaces. If any table name is unqualified, discovery adds the engine default schema as well: public on PostgreSQL, dbo on SQL Server. Two engines do not take a schema list. MySQL treats the connected database as the namespace. MongoDB has collections instead of schemas, so a collection name is always unqualified.

Limits


The extraction engine bounds every job. These are the default values: The Manager rejects a job with more than 10 datasources before it reaches the queue. A result over the size limit fails the job. Fetcher never returns or stores a truncated result. ENGINE_MAX_RESULT_BYTES lowers the size ceiling on the Worker. A request can lower a limit, but it can never raise one above the engine default.

Duplicate jobs


The Manager computes a SHA-256 hash over the whole request — dataRequest and metadata together. It then looks for a job with the same hash created in the last 5 minutes.
  • A match returns the existing job with HTTP 200. No second extraction runs.
  • No match creates a new job and returns HTTP 202.
  • A match that already failed does not block a retry. The Manager creates a new job.
Metadata is part of the hash. Two requests that read the same data under different correlation IDs are different jobs.

Job lifecycle


Before it creates the job, the Manager resolves every datasource name to a connection and opens a real connection to each one. A datasource that fails this test rejects the whole request. Internal datasources configured through environment variables skip the test. A Worker claims a job only while the job is still pending. It then moves the job to processing and starts the extraction. Datasources run in parallel up to the concurrency bound. The run is fail-fast: the first datasource that fails ends the job, and Fetcher stores no partial result. On success the Worker writes the result to object storage and records two values on the job: the result path and the HMAC signature of the result. It then publishes the terminal event.

Next steps


Filters

The ten filter operators and the value shape each one takes.

Datasources

What behaves differently on each of the five database engines.