The dltHub REST API updated yesterday
No OpenAPI endpoint. So I unzipped the client library and diffed it.
TL/DR: dltHub renamed scripts to jobs and runs to job-runs in its platform API and the old trigger endpoint responds with a 404 now. Though there is no public OpenAPI spec (yet), a spec ships inside every dlthub-client wheel. So I diffed the most recent (public) versions and found: between 0.28.6 and 0.28.7 the API flipped.
I trigger dltHub jobs using Snowflake tasks and a Python procedure (who needs an orchestrator when you have async(call ...) and CRON?). Until yesterday it called api.dlthub.com with plain requests. Five Python functions use the dlthub_sdk and are used by a ❄️ Cortex agent - my dltHub job inspector:

Where to find something about that API
The API is used to read and configure and trigger stuff on dltHub, but when you just search for "dltHub REST API" pretty much all the results are about dlt's REST API source 😅
Public record for that API is thin. The docs cover API keys, but (today) contain no endpoint reference. A teaser on the managed-infrastructure marketing page shows POST /v1/jobs/{id}/runs and GET /v1/runs/{id}, which looks nothing like the endpoint my procedure called. That page appears to be a preview sketch...

The wheel is the spec
An OpenAPI document generates the Python client (one _gen/api/api/<tag>/<operation>.py file per endpoint, the URL template hard-coded in _get_kwargs). Hence, the schema is described inside every dlthub-client release, just not as a JSON file (though the kind folks at dltHub most certainly have one and could provide it if asked).
pip index versions dlthub-client # 0.27.0 ... 0.28.7
pip download dlthub-client==0.27.0 --no-deps -d old
pip download dlthub-client==0.28.7 --no-deps -d new
python -c "import zipfile,glob; zipfile.ZipFile(glob.glob('new/*.whl')[0]).extractall('new/x')"
grep -rhoE '"/?api/v[0-9]/[^" ]+' new/x | sort -uDownloaded packages get unpacked in a temp directory just to look at what's inside. Diffing and bisecting the lists of paths between 0.27.0 and 0.28.7 revealed:
| Version | Package layout | Trigger path |
|---|---|---|
| 0.27.0, 0.27.9, 0.28.0, 0.28.3 | dlt_runtime only | /scripts/trigger |
| 0.28.4, 0.28.5, 0.28.6 | dlthub_sdk appears | /scripts/trigger (still /runs, still dataplane-access-token) |
| 0.28.7 | dlthub_sdk | /jobs/-trigger |
The friendly dlthub_sdk namespace showed up in 0.28.4, yet the HTTP surface underneath it remained stable until 0.28.7. The server may well have moved before or after the wheel did, but since yesterday (2026-10-07) the old endpoints answer 404.
What changed
The renames are mechanical (scripts became jobs, runs became job-runs) and action endpoints picked up a dash prefix (/-trigger, /-cancel, /-pause). And payload shapes changed significantly:
| Before (0.28.6) | After (0.28.7) | |
|---|---|---|
| Trigger | POST /workspaces/{ws}/scripts/trigger | POST /workspaces/{ws}/jobs/-trigger |
| Trigger body | job_refs, selectors, dry_run | job_ids (UUIDs), selectors, dry_run |
| Trigger response | {"triggered": [{script_id, run_id, status}]} | {"succeeded": [], "skipped": [], "failed": [{ref, detail}]} with job_run_id per item |
| Fetch a run | GET .../runs/{run_id} | GET .../job-runs/{job_run_id} |
| Run fields | time_started, time_ended, duration, prev_run_id | started_at, finished_at, duration_seconds, previous_job_run_id |
| Data plane token | GET .../dataplane-access-token | POST .../-issue-dataplane-token |
Two of those merit a closer look. The request body no longer has job_refs, so a ref like jobs.__deployment__.my_job needs to be resolved to a UUID first (the SDK pages through GET /jobs and matches on job_ref). And when a job gets skipped over a concurrency limit, it no longer lands in the same list as a started one. My old retry loop checked triggered[0].status == 'skipped_concurrency_limit'... this will never match anything again, so retries might have silently stopped.
The Snowflake twist
Before touching anything, I called the live procedure with a dry run. Expected a 404 Not Found and got it. As I said, the folks over at dltHub are very kind, so they warned me of the change and, hence, I saw it coming 😜 Then came one of the SDK functions, which I didn't really expect to break at all:
ERROR: KeyError: 'date_added'date_added belongs to the old run response (the new one says created_at). On purpose, my ❄️ functions declare packages = ('dlthub-client') unpinned: the SDK is the supported surface and should move with the API. Snowflake, however, resolves that package list exactly once at function creation. Then it's frozen. So my functions ran an old SDK against the new server 🙄
Recreating a function re-resolves the package. Easy enough. The KeyError went away, and the new SDK's Python API drift surfaced:
jobs.get(ref=...)becamejobs.get_by_ref(ref)(gettakes a UUID now)- run field
ended_atbecamefinished_at LogLine.phasebecameLogLine.stage, and the "drop the dependency-install noise" filter has to look atLogLine.source == "setup"instead- the pipeline run's
rows_loadedbecametotal_rows_loaded
One at a time each of those surfaced, because every function swallows its exceptions into an ERROR: ... string for the agent. The agent likes that... I found it annoying 😅 By mapping the new SDK field names back, I kept the output keys the agent sees stable (paused, next_run_at, definition), so the agent spec needed no change.
The new trigger procedure
The interesting part of the changed procedure (full procedure at the bottom if anyone is interested):
def _trigger(workspace, job_ref: str, dry_run: bool, max_retries=3):
for attempt in range(max_retries):
try:
return workspace.jobs.trigger(refs=[job_ref], dry_run=dry_run)[0]
except (dlthub_sdk.errors.TransportError, dlthub_sdk.errors.ServerError) as e:
if attempt == max_retries - 1:
raise Exception(f"dltHub trigger of {job_ref} failed after {max_retries} attempts: {type(e).__name__}: {e}")
time.sleep(2 ** attempt) # 1s, 2sThe rest (polling loop, terminal-state check, etc.) stays as it was. Mostly. The poll target becomes /job-runs/{id} and job_run_id replaces run_id.
Testing it
First everything went in as scratch copies in my SANDBOX and I called those. The dry run came back with succeeded: [{job_id: ..., status: "triggered", job_run_id: null}], the new contract did its thing. For a recent run of the same job, all five functions already using the SDK returned real data (record, run list, job definition, log tail, pipeline trace).
Then the live deploy and a real wait_until_done run of a small job (about two minutes, roughly 14,700 rows). Triggered by the procedure, polled via /job-runs/{id}, status: completed came back after 120 seconds 🎯
And that's the whole trick: the API is whatever the newest wheel claims, and a pip download plus unzip reads it just fine 😎
The full Snowflake procedure to trigger dltHub jobs*
create or replace procedure meta.load.p_trigger_dlthub_pipeline(
job_name varchar
, dry_run boolean default false
, wait_until_done boolean default false
, max_wait_sec integer default 1800 -- only used when wait_until_done; safety valve, bump as needed
)
copy grants
returns variant
language python
runtime_version = '3.13'
artifact_repository = snowflake.snowpark.pypi_shared_repository
packages = ('snowflake-snowpark-python', 'dlthub-client')
handler = 'main'
comment = 'Trigger a dltHub pipeline run via dlthub_sdk; optionally wait synchronously for completion (max 30 min by default), raising if it does not complete'
external_access_integrations = (i_dlthub)
secrets = ('api_key' = meta.integration.se_dlthub_api_key)
as
$$
import _snowflake
import dataclasses
import json
import time
import dlthub_sdk
from dlthub_sdk.domain.jobs import TriggerStatus
WORKSPACE_ID = '<my workspace ID>'
API_BASE_URL = 'https://api.dlthub.com'
def _plain(obj):
# SDK objects carry datetimes and enums: round-trip through json to get a plain variant
return json.loads(json.dumps(obj, default=str))
def _trigger(workspace, job_ref: str, dry_run: bool, max_retries=3):
# retry with backoff on transient transport failures / 5xx - the trigger call to
# api.dlthub.com occasionally hits a ConnectTimeout that used to fail the caller outright.
for attempt in range(max_retries):
try:
return workspace.jobs.trigger(refs=[job_ref], dry_run=dry_run)[0]
except (dlthub_sdk.errors.TransportError, dlthub_sdk.errors.ServerError) as e:
if attempt == max_retries - 1:
raise Exception(f"dltHub trigger of {job_ref} failed after {max_retries} attempts: {type(e).__name__}: {e}")
time.sleep(2 ** attempt) # 1s, 2s
def main(session, job_name: str, dry_run: bool = False, wait_until_done: bool = False, max_wait_sec: int = 1800):
api_key = _snowflake.get_generic_secret_string('api_key')
job_ref = f'jobs.__deployment__.{job_name}'
workspace = dlthub_sdk.connect(token=api_key, base_url=API_BASE_URL).workspaces.get(id=WORKSPACE_ID)
# 1) Trigger the job
result = _trigger(workspace, job_ref, dry_run)
if not wait_until_done or dry_run: # dry runs don't actually execute, so there's nothing to poll
return {
'mode': 'sync' if wait_until_done else 'async',
'job_ref': job_ref,
'dry_run': dry_run,
'response': _plain(dataclasses.asdict(result)),
}
poll_interval = 10 # seconds
elapsed = 0
# job already has an active run (concurrency limit reached) - retry the trigger
# instead of failing outright, until a run is granted or max_wait_sec is used up.
while result.status is TriggerStatus.SKIPPED_CONCURRENCY_LIMIT:
if elapsed >= max_wait_sec:
raise Exception(f"dltHub job '{job_name}' still concurrency-limited after {elapsed}s: {result}")
time.sleep(poll_interval)
elapsed += poll_interval
result = _trigger(workspace, job_ref, dry_run)
if not result.started:
raise Exception(f"dltHub job '{job_name}' was not triggered: {result}")
# 2) Wait until terminal state or timeout (elapsed carries over any time spent
# retrying the trigger above, so max_wait_sec bounds the whole call)
run = workspace.job_runs.get(id=result.job_run_id)
try:
run = run.wait(timeout=max(1, max_wait_sec - elapsed))
except dlthub_sdk.errors.WaitTimeout:
raise Exception(
f"dltHub job '{job_name}' (run_id={result.job_run_id}) did not reach a terminal state within {max_wait_sec}s. "
f"Last known status: {workspace.job_runs.get(id=result.job_run_id).status}"
)
run_info = _plain(run.to_dict())
if run.status.value != 'completed':
raise Exception(
f"dltHub job '{job_name}' (run_id={result.job_run_id}) finished with status '{run.status.value}'. "
f"run_info: {run_info}"
)
return {
'mode': 'sync',
'status': run.status.value,
'run_id': result.job_run_id,
'elapsed_sec': round(elapsed + (run.duration_seconds or 0)),
'run_info': run_info,
}
$$
;License
This is free and unencumbered software released into the public domain. Anyone is free to copy, modify, publish, use, compile, sell, or distribute this software, either in source code form or as a compiled binary, for any purpose, commercial or non-commercial, and by any means.
In jurisdictions that recognize copyright laws, the author or authors of this software dedicate any and all copyright interest in the software to the public domain. We make this dedication for the benefit of the public at large and to the detriment of our heirs and successors. We intend this dedication to be an overt act of relinquishment in perpetuity of all present and future rights to this software under copyright law.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
For more information, please refer to https://unlicense.org/
