M2M Product Sync Guide
This guide explains the recommended way to synchronize the TBC product catalog with your external system (e.g., WooCommerce, ERP, or custom app).
[!IMPORTANT] Use the Search Endpoint (
/products/search), NOT the List Endpoint (/products).The standard
/productsendpoint returns a limited subset of fields and is deprecated for sync purposes. The/products/searchendpoint provides full product data and robust incremental sync capabilities.
Sync Strategy Overview
Section titled “Sync Strategy Overview”1. The Challenge
Section titled “1. The Challenge”Synchronizing a large catalog (200,000+ products) is resource-intensive. A daily “full download” is inefficient and prone to timeouts.
2. The Solution: Incremental Sync
Section titled “2. The Solution: Incremental Sync”Instead of downloading everything every time, you should:
- Initial Sync: Download all products once.
- Incremental Sync: Periodically (e.g., hourly) ask only for products that have changed since your last check.
We support this natively using the updatedAfter parameter.
Prerequisites
Section titled “Prerequisites”Ensure you have your Client ID and Client Secret and can successfully authenticate to get a Bearer token. See Authentication Guide
1. Initial Sync (Full Download)
Section titled “1. Initial Sync (Full Download)”To perform your first full download, you will query for all products “updated after the beginning of time”.
Endpoint: GET /products/search
| Parameter | Value | Description |
|---|---|---|
updatedAfter | 1970-01-01T00:00:00Z | Requests the whole non-archived catalogue, sorted stably by time (see the guarantees below). |
size | 50 | Number of products per page (max recommended). |
syncToken | (Dynamic) | The cursor for the next page of results. |
How it works
Section titled “How it works”- Request page 1 with
updatedAfter=1970-01-01T00:00:00Z. - Process the products in
data. - Check if
syncTokenis present in the response. - If yes, request the next page passing
syncToken. - Repeat until
syncTokenis null.
What a completed full pass guarantees (and what it doesn’t)
Section titled “What a completed full pass guarantees (and what it doesn’t)”A full pass reads a live index while it changes — there is no point-in-time snapshot
behind syncToken paging. “syncToken is null” means the cursor reached the end of the
result set, not that you hold a coherent copy of the catalogue as of a single moment:
- Duplicates are expected. Results are ordered by
updatedAt, so a product edited while your pass is running moves ahead of the cursor and can be returned again, carrying its newer data (whether it is depends on when the edit becomes visible in the index relative to your pass). This is harmless if you upsert idempotently byid— the record is simply written twice. - Omissions are possible, and silent. The index is updated asynchronously after a product is saved. A product whose indexing completes only after your cursor has passed its position is missed by that pass, with no error — the same indexing-lag gap that motivates the incremental overlap window. A missed product is normally picked up by a later incremental sync or full pass once it is indexed — though those later passes read the same live index and carry the same caveat. A product that never indexes at all (a pipeline failure on our side) will not appear in any pass — that is a signal to report to us, not something client-side logic can recover.
- Archived products silently leave the feed. Sync results exclude archived products by default, so archiving looks the same as the product never having existed — no pass will tell you to remove it downstream. (Better archive propagation is tracked as a separate improvement.)
- Never delete by set-difference from a single pass. Because one pass can miss records, “in my store but not in this pass” does not mean “removed from the CRM”. Deleting on that basis will destroy live products.
Treat the feed as convergent, not transactional: idempotent upserts by id, frequent
incremental syncs with the overlap window, and a periodic full pass as a reconciliation
sweep together keep your copy converging on the CRM’s non-archived catalogue and keep any
drift short-lived — but no single pass, and no fixed number of passes, is guaranteed
complete. If a specific product still hasn’t arrived after a reconciliation sweep, report
it to us rather than working around it.
2. Incremental Sync (Ongoing Updates)
Section titled “2. Incremental Sync (Ongoing Updates)”Once you have a full copy, you only need to fetch updates.
Endpoint: GET /products/search
| Parameter | Value | Description |
|---|---|---|
updatedAfter | {Last_Success_Timestamp} | Example: 2024-03-20T10:00:00Z |
size | 50 | Batch size. |
syncToken | (Dynamic) | Pagination cursor. |
- Store the timestamp of your last successful sync (e.g., in a database/option).
- Request
updatedAfter={stored_timestamp}. - Process updates (create new items, update existing ones).
- Loop through pages using
syncTokenuntil finished. - Crucial: Update your stored timestamp to the current time only after a successful batch.
Implementation: Using WPGetAPI (WordPress)
Section titled “Implementation: Using WPGetAPI (WordPress)”If you use the WPGetAPI plugin for WooCommerce, follow these settings.
Step 1: Endpoint Configuration
Section titled “Step 1: Endpoint Configuration”In WPGetAPI settings, add a new endpoint:
- Endpoint ID:
get_products_search - Method:
GET - Endpoint:
products/search - Query String:
(Note: The PHP code below will dynamically overwrite these query args)size=50updatedAfter=1970-01-01T00:00:00Z
Step 2: Add PHP Filter for Pagination
Section titled “Step 2: Add PHP Filter for Pagination”Accessing the search endpoint requires a different pagination logic (syncToken instead of page or lastKey). Add this to your theme’s functions.php or a code snippet plugin.
/** * Handle Pagination for TBC Product Search Endpoint * * This sets up the correct pagination parameters for the WPGetAPI plugin * specifically for the /products/search endpoint. * * IMPORTANT: The filter callbacks receive plain string IDs, NOT objects. * See https://wpgetapi.com/docs/pagination for official documentation. */
// WARNING: Replace 'tbc-crm-api' with your actual API Unique ID from WPGetAPI settings// (visible in the Template Tag on your API's settings page)// If they don't match, pagination will fail silently!
// 1. Tell the plugin we are using a custom token, NOT a page numberadd_filter('wpgetapi_api_to_posts_pagination_type', 'tbc_search_pagination_type', 10, 3);
function tbc_search_pagination_type($type, $api_id, $endpoint_id) { if ($api_id === 'tbc-crm-api' && $endpoint_id === 'get_products_search') { return 'syncToken'; // The query param name expected by the API } return $type;}
// 2. Extract the token from the API response dataadd_filter('wpgetapi_api_to_posts_pagination_next_value', 'tbc_search_pagination_token', 10, 4);
function tbc_search_pagination_token($next_value, $data, $api_id, $endpoint_id) { if ($api_id === 'tbc-crm-api' && $endpoint_id === 'get_products_search') { // $data is the API response body (2nd param per WPGetAPI docs) if (is_string($data)) { $data = json_decode($data, true); }
// The API returns 'syncToken' at the root level $token = $data['syncToken'] ?? null; // CRITICAL: URL encode to prevent + becoming space return $token ? urlencode($token) : null; } return $next_value;}Step 3: Configure “API to Posts” Mapping (Crucial)
Section titled “Step 3: Configure “API to Posts” Mapping (Crucial)”If you are using the API to Posts importer, you must tell it where the products are located in the response.
- Go to API to Posts settings for your import.
- Locate the “Results / Items Key” or “Data Location” setting.
- Set this to:
data- Explanation: The Search API returns products inside a
"data": [...]array, whereas the old endpoint used"products". If you leave this blank or default, the importer will fail with an error likeUndefined array key "products".
- Explanation: The Search API returns products inside a
Step 4: Running the “Updates Only” Sync
Section titled “Step 4: Running the “Updates Only” Sync”To run an incremental sync, you cannot just click “Run” in the plugin settings (which uses the hardcoded 1970 date). You must trigger it programmatically or use a separate endpoint configuration for daily updates vs initial backfill.
For a completely automated “Set and Forget” solution using WPGetAPI, we recommend:
- Config A (Backfill): Hardcoded
updatedAfter=1970.... Run this ONCE. - Config B (Updates): Use a custom PHP function to call the API with the dynamic date.
PHP Integration Example
Section titled “PHP Integration Example”This is a robust example of how to run the daily sync via code, bypassing the plugin UI manual trigger.
function tbc_run_daily_sync() { // 1. Get last sync time (default to epoch if never run) $last_sync_time = get_option('tbc_last_sync_time', '1970-01-01T00:00:00Z');
// 2. Capture the boundary BEFORE fetching, then rewind it by an overlap // window. Both halves matter: // - capturing before the pass avoids losing records updated while the // pass is running // - the overlap absorbs INDEXING LAG. Search is served from a separate // index kept in sync asynchronously, so a product saved just before // the boundary may not be searchable until after this pass has // finished. Without the rewind it is absent now and excluded next // time, which loses it permanently. // Re-fetching a handful of already-seen products is harmless as long as // you upsert by `id` (see below). Losing one is not. $overlap_seconds = 900; // 15 minutes $boundary = gmdate('Y-m-d\TH:i:s\Z', time() - $overlap_seconds);
// 3. Prepare arguments $token = null; // Start with no token $ok = true; // Track sync success
do { $query_args = array( 'size' => 50, 'updatedAfter' => $last_sync_time );
if ($token) { $query_args['syncToken'] = $token; }
// Call WPGetAPI programmatically // Format: wpgetapi_endpoint( $api_id, $endpoint_id, $args ) $api = wpgetapi_endpoint('tbc-crm-api', 'get_products_search', array('query_variables' => $query_args));
// Transport failure: WPGetAPI hands back a WP_Error object. It is not a // JSON string, so json_decode() below would throw, and it is not // array-accessible either. Handle it first. if (is_wp_error($api)) { error_log('TBC sync incomplete: ' . $api->get_error_message()); $ok = false; break; // abort WITHOUT committing the checkpoint }
// WPGetAPI returns EITHER a decoded array OR the raw JSON body, depending // on the endpoint's "Format results" setting. Handle both, or one // configuration makes every response look malformed and the sync aborts // forever without ever committing a checkpoint. $data = is_array($api) ? $api : json_decode($api, true);
// A successful response ALWAYS carries a `data` array; an empty final page // is `data: []`. No error response carries `data` at all. Checking for the // presence of success is robust to every error shape, including failures // that carry no machine-readable code. if (!is_array($data) || !isset($data['data']) || !is_array($data['data'])) { // $data may be null (json_decode failure) or a scalar; neither is // array-accessible, so only read offsets once we know it is an array. $detail = is_array($data) ? ($data['error'] ?? $data['code'] ?? $data['message'] ?? 'unknown') : 'malformed response'; error_log('TBC sync incomplete: ' . $detail); $ok = false; break; // abort WITHOUT committing the checkpoint }
if (empty($data['data'])) { break; // genuine end of the result set }
// 4. Process Products foreach ($data['data'] as $product) { // Your logic to save product to WooCommerce // update_woocommerce_product($product); }
// 5. Get next token $token = $data['syncToken'] ?? null;
} while ($token);
// 6. Update last sync time ONLY on a clean pass if ($ok) { update_option('tbc_last_sync_time', $boundary); }}Response Field Reference
Section titled “Response Field Reference”The /products/search endpoint projects every product through an external field
allowlist before it leaves the API. This is the complete list — nothing else
comes back, regardless of what the underlying product record contains:
id: Internal UUID (immutable — see Idempotent Upsert by ID below)productCode: The SKUname,description: Naming/description fieldsbrandName,brandSlug,range: Brand and rangecategory,department,subDepartment,stockType,catalogue: ClassificationdisplayOnWeb: Whether the product should be published/shown on your site (e.g. WooCommerce published vs. draft). If your API client is bound to a web channel, this is always an explicittrue/falsereflecting that channel’s visibility. If your client is unbound, it reflects the product’s stored visibility flag, which may be omitted when no flag is stored — treat a missing value as not visible.notes: Short public-facing description, used as a public short description, not private staff content. Still labelled “Internal Notes” in the TBC CRM UI pending the #445 Task 2 relabel to “Short Description”. A separateinternalNotefield holds genuinely-private staff notes and is never exposed via this feed.colour: Product colour/finishprice: Trade/B2B priceretailPrice: Recommended retail price (RRP)supplierCost: Wholesale/cost-basis pricebarcodelength,height,width,weight,weightUnit: DimensionscreatedAt: Timestamp the record was createdimageUrl,images(array of{url, name}): Product imagerydocuments(array of{url, name}): Product documents (spec sheets, manuals)erpCode,altSupplierCode,altSupplierCode2: Supplier/ERP identifiers —erpCodeis also whatq=search matches against for identifier lookupswebHierarchy: Website navigation/category tree path, e.g.Bathrooms;Accessories;Heated Towel Railstype,subType,subSubType: Product classification treeretailRange: Product range/collection name shown on your site — distinct fromrange, which is frequently empty for the same productsrelatedProducts(array): A projected view of the product’s related-items relationships —{productId, productCode?, erpCode?, relationType?, qty?, source?}.productCode/erpCodeare inline display fields present only onGET /products/{productId}(the API enriches them before projecting);GET /products/searchentries carry onlyproductId/relationType/qty/source.source: 'reverse'marks a server-derived “used by” relation, as opposed to a forward “Required Items” entry.note,name,imageUrl, andarchivedare never present; the separatesparesarray is never exposed at all.
If you use the fields= parameter to request a subset of the response,
requesting a field outside this list is narrowing-only: the denied field
is silently omitted, never added, and no error is raised.
Example Response:
{ "data": [ { "id": "123-abc", "productCode": "SKU-999", "description": "Luxury Mixer", "price": 5400.00, ... } ], "pagination": { "total": 1500, ... }, "syncToken": "eyJ..."}Error Handling Contract
Section titled “Error Handling Contract”The endpoint may return errors during incremental sync. Understanding the error contract is critical to avoid silent data loss.
The Rule: Check for Success Signal
Section titled “The Rule: Check for Success Signal”A successful response ALWAYS contains a data array. An empty final page returns data: []. No error response contains data at all — not even when the error has a machine-readable code.
This is the robust, portable check that works across all error shapes:
// A successful response ALWAYS carries a `data` array; an empty final page// is `data: []`. No error response carries `data` at all. Checking for the// presence of success is robust to every error shape, including failures// that carry no machine-readable code.if (!is_array($data) || !isset($data['data']) || !is_array($data['data'])) { // Read offsets only once you know $data IS an array: a transport failure or // a json_decode() failure leaves it null or scalar, and indexing that // fatals — killing the request before the checkpoint guard below can run, // which is the exact failure this block exists to prevent. $detail = is_array($data) ? ($data['error'] ?? $data['code'] ?? $data['message'] ?? 'unknown') : 'malformed response'; error_log('TBC sync incomplete: ' . $detail); $ok = false; break; // abort WITHOUT committing the checkpoint}
if (empty($data['data'])) { break; // genuine end of the result set}Do NOT commit your checkpoint and do NOT advance your cursor if data is absent.
How to retry depends on why it failed, and the cases are opposites:
- Transient (
SEARCH_INCOMPLETE,SEARCH_TIMEOUT,IDENTIFIER_LOOKUP_INCOMPLETE, or any 5xx with no code) — retry after a delay, keeping your cursor. See Retrying safely. INVALID_SYNC_TOKEN— do not retry the same request. That cursor is permanently unusable and resending it will fail identically forever. Discard it and restart the pass from your last committed checkpoint.- Auth failures — a 401 means your token expired: fetch a new one and retry. A 403 means your client is not permitted this scope; retrying will never succeed, so stop and alert a human.
A failure with no code is NOT automatically transient. Authentication and authorisation
rejections are produced by the API gateway before our handler runs, so they carry message only —
no error, and no data. Retrying a 403 forever accomplishes nothing.
Classify by HTTP status first, and use error only to refine within a status:
| status | meaning | action |
|---|---|---|
| 401 | token expired | re-authenticate, retry once |
| 403 | scope not permitted | stop, alert — never retry |
| 429 | rate limited by the gateway or WAF | retry with backoff — this is the one retryable 4xx |
| 4xx other | malformed request (incl. INVALID_SYNC_TOKEN) | fix or discard cursor; do not blind-retry |
| 5xx / 504 | transient backend failure | retry with backoff |
Note the 429 row: it is a 4xx but it is transient. Treating “any 4xx” as fatal stops a sync that would have succeeded a few seconds later.
If your HTTP client does not expose the status code — some WordPress plugin wrappers do not — then treat an uncoded failure as retryable but strictly bounded: cap attempts (three is plenty), then abort the run and surface it. Never let an uncoded failure retry indefinitely, and never let it commit a checkpoint.
Error Codes (Informational)
Section titled “Error Codes (Informational)”Once you detect failure via the absence of data, the response may contain one of these machine-readable codes to guide your retry strategy:
error: SEARCH_INCOMPLETE(503) — the search could not be proven to have run to completion (it timed out, or a shard failed), so the page you would have received might be missing rows. We return an error rather than the partial page precisely so you never mistake a truncated page for the end of the catalogue. Retry with exponential backoff.error: IDENTIFIER_LOOKUP_INCOMPLETE(503) — Identifier lookup did not complete. Retry with backoff.error: SEARCH_TIMEOUT(504) — the query exceeded its time budget. Retry with backoff, and shrinksize— see Retrying safely.error: INVALID_SYNC_TOKEN(400) — Your cursor is stale or corrupt. Discard it and restart from your last committed checkpoint.
Note: This list is not exhaustive, and some failures carry no code at all. The data check is
what tells you a request FAILED; the HTTP status is what tells you whether to retry. Use error
only to refine within a status, and for logging.
Retrying safely
Section titled “Retrying safely”Exponential backoff alone is not enough when the same page keeps failing. SEARCH_INCOMPLETE
usually means that particular request was too expensive to complete inside the server’s time
budget — retrying it unchanged will most likely fail the same way, however long you wait.
- Cap attempts per page. Three is plenty. After that, abort the run and surface it on your side rather than looping.
- Shrink the page before retrying. If a page fails twice, halve
sizeand try again. A smaller page is materially more likely to complete, and it is the single most effective thing you can change. This matters most on the initial backfill, whereupdatedAfteris the epoch and the result set is the whole catalogue. - Never advance the checkpoint on an aborted run. A partially-completed pass must be retried from the same boundary, not resumed from a later one.
- Leave
includeAggregationsoff unless you actually consume the facet counts. It adds real work to every page.
If you have backed off, halved the page, and it still fails, that is worth telling us about —
include the requestId from the response body.
Why the overlap window matters
Section titled “Why the overlap window matters”Search results come from an index that is updated asynchronously after a product is saved. Normally that catch-up takes seconds. It is not guaranteed to.
That gives incremental sync a gap that a naive checkpoint cannot see:
14:59:50 product P is saved in the CRM15:00:00 your pass starts, checkpoint boundary = 15:00:0015:00:02 your pass queries — P is not in the index yet, so it is not returned15:00:30 P finishes indexing next run asks for updatedAt > 15:00:00 P's updatedAt is 14:59:50, so it is never returned againRewinding the committed boundary by an overlap window closes it: the next pass asks from
15:00:00 - 15 minutes, P falls inside that, and you receive it. You will also re-receive
everything else from those 15 minutes, which costs nothing if your writes are idempotent.
The overlap is not a complete backstop. It covers ordinary lag. If the indexing pipeline
stalls for longer than your window, records can still fall through. If exact completeness
matters to you, run a periodic full pass (updatedAfter=1970-01-01T00:00:00Z) on a slower
cadence — weekly is usually enough — as a reconciliation sweep alongside the frequent
incremental one. Note that the full pass pages the same live index, so it has its own
mid-pass version of this gap — see
What a completed full pass guarantees.
Tell us if you see a product that never arrives; that is a signal on our side,
not yours.
Best Practice: Idempotent Upsert by ID
Section titled “Best Practice: Idempotent Upsert by ID”Since the boundary race condition is inherent to any “checkpoint at the end” design, always upsert products by their id field (not by position, and not by SKU). This way, if you fetch the same product twice due to a retry or checkpoint race, your record is simply updated again — no duplicates.
Use id, not productCode. id is the CRM’s immutable primary key. productCode is a
business identifier that can be edited — and when it is, a SKU-keyed lookup finds nothing and
creates a second product instead of renaming the first. You then have a duplicate that no
subsequent sync will ever reconcile, because the original is now unreachable by its new code.
Store the CRM id alongside your product and look up by that, falling back to SKU only for
records created before you started storing it:
foreach ($data['data'] as $product) { // Look up by the immutable CRM id first. $existing = get_posts(array( 'post_type' => 'product', 'meta_key' => '_tbc_product_id', 'meta_value' => $product['id'], 'posts_per_page' => 1, 'post_status' => 'any', // get_posts() defaults to 'publish' 'fields' => 'ids', )); $product_id = $existing ? $existing[0] : 0;
// Fallback for products imported before _tbc_product_id was stored. // Remove this branch once your catalogue is fully backfilled. if (!$product_id) { $product_id = wc_get_product_id_by_sku($product['productCode']); }
$wc_product = $product_id ? wc_get_product($product_id) : new WC_Product_Simple();
$wc_product->set_name($product['description']); $wc_product->set_sku($product['productCode']); // may CHANGE between syncs // ... set other fields ... $wc_product->save();
// Record the CRM id so the next sync matches on identity, not on SKU. update_post_meta($wc_product->get_id(), '_tbc_product_id', $product['id']);}Troubleshooting
Section titled “Troubleshooting”Pagination Loop (Same Products Repeatedly)
Section titled “Pagination Loop (Same Products Repeatedly)”Symptom: WPGetAPI keeps importing the same products on every page.
Root Causes (in order of likelihood):
- Wrong filter parameters: The WPGetAPI pagination filters pass
$api_idand$endpoint_idas plain strings, and the API response as$data(the 2nd parameter). If your code treats these as objects (e.g.,$api->id,$api->response_body), the syncToken will never be extracted. See the official WPGetAPI pagination docs. - API ID mismatch: If the API Unique ID in your PHP code doesn’t match your WPGetAPI settings, the filters are silently skipped.
- Missing URL encoding: The
syncTokencontains+characters (Base64) that become spaces withouturlencode().
Solution: Ensure your code matches Step 2 above — use $data (2nd param) for response data, $api_id/$endpoint_id (3rd/4th params) as strings, and wrap with urlencode().
Verification: Our E2E tests confirm the API pagination works correctly:
- Page 1 returns a valid
syncToken - Page 2 with
syncTokenreturns different products - No loop detected when filters are correctly implemented
Debugging Checklist:
- Verify API Unique ID: Check the Template Tag on your API settings page — the first quoted string is your API ID. It must match your PHP code exactly.
- Check WPGetAPI debug logs for the page 2 request URL
- Verify
syncToken=parameter is present in the URL - Check for HTTP 200 response (not 400)
Summary
Section titled “Summary”- Use
/products/searchfor all product syncing. - Start with
updatedAfter=1970-01-01T00:00:00Zfor the initial backfill. - Switch to incremental syncs with a stored timestamp for ongoing updates.
- Always use
urlencode()when passing thesyncToken. - The
dataarray contains the products, andsyncTokenis your pagination cursor.