Service 23

CRM integration that keeps one system right and the rest in step

Sync between your CRM and the systems around it, with one owner per field, idempotent upserts, a conflict log, and a reconciliation job that proves both sides still agree.

On this page
  1. What CRM integration actually is
  2. Who it is for, and who it is not for
  3. What we actually build
  4. How it works technically
  5. Field custody, the rule that prevents most sync disasters
  6. The build process stage by stage
  7. What you get at handover
  8. Where CRM integrations go wrong
  9. What it costs to run once live
  10. How to tell whether you need this
  11. How to start

What CRM integration actually is

The short answer

CRM integration is the work of making your CRM agree with the other systems that hold the same facts: the product database, the billing system, the support desk, the marketing platform. Done properly it is not a pipe between two apps. It is a written decision about which system owns each field, a sync that applies changes idempotently in that direction, a log of every conflict, and a scheduled reconciliation that checks both sides still match. The connector is the easy part. The ownership decisions are the project.

Most CRMs die the same way. Three systems write to the same field, nobody agreed who should, and after a few months the sales team stops trusting the record and keeps the real state in a spreadsheet. At that point the CRM is a place data goes to be forgotten, and no amount of new tooling fixes it, because the problem is not connectivity. It is that nothing decides what is true.

So the first deliverable is never code. It is a field-level map naming one writer per field, with the exceptions listed and a survivorship rule attached to each. Once that document exists, the engineering is mostly mechanical.

  • One writerEvery field names exactly one system permitted to write it. Everything else reads. Fields with two writers are the source of most sync incidents.
  • Natural keysRecords match on a normalised real-world key, a lowercased email or a cleaned company domain, not on a record id that only one system understands.
  • Conflicts loggedWhen a survivorship rule fires, both values, the rule and the winner are recorded. That log is how you discover the rule is wrong.
  • Reconciled nightlyA scheduled job compares both systems field by field and opens a ticket on drift, because a sync that has silently stopped looks exactly like a sync with nothing to do.

Who it is for, and who it is not for

This is for companies where the CRM is one of several systems holding customer state, and where the same fact is currently entered by hand in more than one place. If your CRM is the only system and a native connector covers your two other tools, buy the connector.

Your situationVerdictWhy
A CRM plus billing, product and support all holding customer stateGood fitFour systems means twelve possible pairings. Ownership has to be declared or it will be improvised.
Someone re-keys the same field into two systems every weekGood fitThat is a mapping and a trigger, and re-keying is where the divergence starts.
Native connector already covers your two tools and nobody complainsNot a fitBuy the connector. Custom work here is expensive and adds a thing to maintain.
You are about to change CRMWaitBuild the integration against the CRM you are keeping, not the one you are leaving.
Records are already duplicated and nobody knows the real countFix firstSyncing on top of duplicates multiplies them. Dedupe is a separate project and it comes first.
You want the integration to clean bad data on the way throughCarefulNormalisation yes, invention no. A sync that guesses missing values makes the data worse and harder to audit.
The fifth row is the one that most often turns a six-week build into a four-month one.

Three things that have to exist before the build

  1. An agreed natural key per object. Usually a lowercased email for people and a normalised domain for companies, with the known exceptions written down: shared inboxes, subsidiaries, personal addresses on business accounts.
  2. API access at a tier with enough call allowance, plus somebody who can raise it. Integrations that work in testing and fail in production are usually hitting a ceiling nobody checked.
  3. A named owner for the CRM schema. If anyone in the company can add a required field or a validation rule on a Tuesday afternoon, the integration will break on a Tuesday afternoon.

Who should not buy this

  • Teams whose real problem is CRM adoption. If sales will not update the record, syncing more data into it changes nothing.
  • Anyone hoping the integration will settle an argument about which department owns the customer. It will surface the argument, not resolve it.
  • Companies mid-way through a data migration. Finish the migration, then integrate the result.
  • Teams who need reporting rather than sync. That is a warehouse job, and building it as an integration makes it slower and more fragile.

What we actually build

Six components. Most of this build contains no model at all, which is worth saying plainly on a page hosted by an AI automation studio.

The field custody manifest

A version-controlled file listing every object and field in scope with its writer, its readers, its transform, its allowed values and its conflict rule. It is readable by an operations lead, it is the artefact you argue about, and it is what the sync compiles from. When it changes, the change goes through review like any other code.

The identity and dedupe layer

Normalisation and matching before anything writes. Emails lowercased and trimmed, domains stripped of subdomains and free-mail providers flagged, company names normalised for matching but never overwritten. Deterministic rules first. Fuzzy matching only with a confidence threshold and a review queue, because an automatic merge of two real customers is one of the few integration errors that is genuinely hard to undo.

The sync workers

One worker per direction, driven by webhooks where the platform offers them and by a polled watermark where it does not. Each write is an upsert on the natural key, carries an idempotency key, and is tagged with its source so that the echo event coming back does not start a loop. Batched where the API rewards batching, serialised per record so two updates to the same contact cannot race.

The conflict log and the review queue

Every time a survivorship rule fires, the system records both values, the rule applied, the winner and the source events. Anything the rules cannot settle goes to a queue a human clears. Teams often discover their real ownership policy by reading a fortnight of that log, which is a better outcome than the workshop that was supposed to produce it.

The reconciliation job

Nightly, it walks both systems and compares the fields in scope. Drift opens a ticket with the record ids attached. This is the component that catches the failure everyone misses: a sync that stopped days ago and produced no errors, because zero events processed looks identical to a quiet weekend.

The backfill, run separately

History moves through its own path, with its own rate limit budget, its own batch size and a resumable cursor. Running a backfill through the live path is how teams consume a day's API allowance by lunchtime.

Native connector or iPaaSAn integration you own
Field-level ownershipObject level, sometimes direction onlyPer field, declared in a reviewable file
Custom objects and fieldsOften unsupported or extra-costWhatever your schema actually contains
Conflict handlingLast write wins, usually invisibleStated rule, logged, queue for the rest
Backfill controlLimited, and it shares your live quotaSeparate path, own budget, resumable
Time to first valueDaysWeeks
Who fixes it at 2amA support ticketYour on-call, with logs and a replay
Use the connector until it hurts

Start with the native connector and note every time it forces a compromise: a field it cannot map, a direction it cannot express, a conflict it resolved wrongly. After a month you have a specification written from evidence rather than from a whiteboard, and you may find the connector was fine.

How it works technically

Events in, identity resolved, custody applied, write attempted, result verified, drift reconciled. The parts that break in production are identity, loops and rate limits, in that order.

Triggers, watermarks and the ordering problem

Webhooks are cheaper and faster, and they are not reliable enough to be the only path: they can be delayed, duplicated or dropped. Every webhook-driven sync also needs a polled sweep on a watermark to catch what was missed. Advance the watermark only after the write has been confirmed, never before, because the crash between those two lines is the classic silent data loss. Ordering matters too, since two updates to the same record arriving out of order will leave the older value in place. Serialise per record key rather than globally. The trade-offs are covered in webhooks vs polling.

Writes, echoes and loops

A two-way sync without loop suppression is a machine for generating infinite updates. System A writes, the webhook fires, system B writes back, A hears its own change and writes again. Three defences, and you want all three: tag every write with its source and ignore inbound events carrying your own tag, compare the new value with the current value and skip no-op writes, and cap updates per record per hour so a loop trips a circuit breaker instead of running all night.

Rate limits and the CRM's own automation

CRM APIs meter calls, and the allowance depends on your product tier, so read your own limits page rather than any blog. Budget calls per record before you build: a lookup, a dedupe check and a write is three calls, and a fifty-thousand record backfill is a hundred and fifty thousand of them. A write into a CRM can also fire the CRM's own workflows. A backfill touching a field watched by an email automation is how a company sends a very large number of unintended emails in one afternoon.

Field custody manifest, the file the whole sync compiles fromyaml
# contact.custody.yaml
# One writer per field. Everything else reads. Direction is per field, never per object.
object: contact
natural_key: email_lower          # the dedupe key, deliberately not the CRM record id

sync:
  trigger: webhook                # poll only where webhooks do not exist
  poll_fallback_minutes: 15
  batch_max: 200
  loop_suppression: source_tag    # ignore an inbound event this system caused
  idempotency_key: "sha256(object + natural_key + field + new_value + source_event_id)"

fields:
  email:
    writer: product_db
    readers: [crm, billing, support]
    transform: lower_trim
    pii: true
  lifecycle_stage:
    writer: crm
    readers: [product_db, marketing]
    allowed_values: [lead, mql, sql, customer, churned]
    on_invalid: reject_and_log    # never coerce a value you do not recognise
  plan_tier:
    writer: billing
    readers: [crm, support]
  last_active_at:
    writer: product_db
    readers: [crm]
    write_if: "new_value > current_value"   # monotonic, never moves backwards
  owner_id:
    writer: crm
    readers: [support]
    survivorship: newest_non_empty
    on_conflict: log_both_values

deletes:
  strategy: tombstone             # propagate a delete, never a silent no-op
  merges: follow_surviving_id     # CRM merges must repoint foreign keys

dead_letter:
  queue: crm_sync_dlq
  retry: { attempts: 5, backoff: exponential, jitter: true }
  alert_after: 10

reconciliation:
  schedule: "daily 02:00"
  compare: [email, lifecycle_stage, plan_tier, owner_id]
  tolerance: 0
  on_drift: open_ticket

Two lines in that file are worth more than the rest combined. The write_if guard on last_active_at makes the field monotonic, so an out-of-order event cannot move a timestamp backwards. The on_invalid: reject_and_log on the picklist means an unrecognised value is refused loudly rather than coerced into the nearest match, which is how bad data enters a CRM and stays there.

Terms used on this page
Upsert
A write that creates the record when the key matches nothing and updates it when it does. Upserts are only safe against a natural key. Upserting on a system-specific id creates the same person twice.
Natural key
The field or combination that identifies a real entity across systems, such as a lowercased email or a normalised company domain. Weak natural keys are the direct cause of duplicate records.
Survivorship rule
The stated rule deciding which value wins when two systems have written the same field. Common rules are newest timestamp wins, non-empty wins, and human edit beats machine write.
Watermark
The stored timestamp or cursor recording how far a sync has read, so the next run resumes instead of rescanning. A watermark advanced before the write is confirmed causes silent data loss.
Tombstone
A record marking that something was deleted, kept so the deletion can propagate to other systems. Without tombstones, deletes vanish and a deleted contact reappears on the next sync.

Field custody, the rule that prevents most sync disasters

Almost every integration incident we have been called into traces back to a field with two writers and no stated rule. The model below is not sophisticated. It is just written down, which is the part everybody skips.

Framework

The ChatGPTalker Field Custody Model

Five rules covering who may write what, and what happens when the answer is contested. Applied field by field, in a file, before any connector is configured.

01
One writer per field

Every field names exactly one system permitted to write it. Everything else reads. A field with two writers is not a design, it is a bug with a date on it.

02
Direction is per field, not per object

The same contact can have the CRM owning stage and owner, billing owning plan tier, and the product database owning last active date. Object-level two-way sync is where the trouble starts.

03
Exceptions get an explicit survivorship rule

Where two systems genuinely must write, state the rule: newest wins, non-empty wins, or human edit wins. Put it in the manifest rather than in the memory of whoever configured it.

04
Conflicts are logged, never silently resolved

Record both values, the rule applied, the winner and the source events. A fortnight of that log tells you more about your real data policy than any workshop will.

05
Every field carries an expiry review

Fields accrete. Quarterly, flag any field nothing has read in ninety days. An unused field still costs a mapping, a sync call and a possible conflict, and it still holds personal data you must delete on request.

Two-way sync is a phrase, not a design

When a vendor says two-way sync, ask which system wins on a specific contested field, and what happens to the losing value. If the answer is last write wins, the design is that whichever system was slowest to fire is right, which is not a policy anyone would choose deliberately.

The build process stage by stage

Six stages. The first is a data audit, and it is the stage clients most often want to skip and most often thank us for afterwards.

  1. Field-level data auditWeek 1

    Every object and field in scope profiled: fill rate, distinct values, duplicate rate on the proposed natural key, and which system last wrote each field. This is where you find out that a third of your company records share four generic domains.

  2. The custody manifestWeek 1 to 2

    The map gets written and argued about with the people who own each system. Contested fields are the useful part of the meeting. The output is a file, in the repository, with a name attached to every decision.

  3. Identity and dedupeWeek 2 to 4

    Normalisation, matching rules, a duplicate report and, where needed, a merge plan run by humans with the machine proposing. Nothing syncs until the duplicate rate on the natural key is understood and accepted.

  4. One field, end to endWeek 3 to 5

    A single low-risk field synced through the full path: trigger, worker, idempotent write, verification, conflict log, reconciliation. Every architectural mistake shows up here, cheaply, on a field nobody would miss.

  5. Widen the scope, then backfillWeek 5 to 8

    Remaining fields added in batches, each with its custody entry. The backfill runs last, on its own path, in off-peak windows, with the CRM's own workflow automations reviewed and paused where a bulk write would trigger them.

  6. Reconciliation, alerting and handoverWeek 8 to 10

    The nightly comparison goes live, the dead letter queue gets an owner, and heartbeat alerts fire when a sync produces no events for longer than it should. Then documentation and a walkthrough with the person who will own it.

Review the CRM automations before any backfill

A CRM write can fire workflows, notifications and email sequences that were designed for human activity. Before a bulk write, list every automation watching the fields you are about to touch, and pause them deliberately. This is the single most expensive mistake in this category and it takes twenty minutes to prevent.

What you get at handover

The manifest, the workers, the queues and the reconciliation all run in your infrastructure under your credentials. The test of a real handover is whether your team can add a new field to the sync without calling us.

Handover acceptance checklist
0 of 6 done

We also hand over the document nobody asks for: what the integration deliberately does not do, and why. Six months later that page prevents an argument about whether something was forgotten or decided.

Where CRM integrations go wrong

Six failure modes. Every one of them is recoverable if you find it in week one and unpleasant if you find it in month six.

The ping-pong loop

Two-way sync with no loop suppression. A writes, B hears it and writes back, A hears its own echo. The record updates forever, the audit trail fills with machine edits, and the API allowance disappears. Source tagging, no-op comparison and a per-record update ceiling are the three defences.

Duplicate explosion from a weak key

Syncing on a system-specific record id, or on an email address for a population that shares inboxes, produces a second copy of everyone. The copies then sync onward. Profile the duplicate rate on your proposed natural key before writing a line of sync code.

The backfill that eats the day's quota

History pushed through the live path consumes the API allowance, live events queue up behind it, and by afternoon the sync is hours behind with no error to show for it. Separate path, separate budget, off-peak window, resumable cursor.

Deletes and merges that go nowhere

Most integrations sync creates and updates and quietly ignore deletions and merges. CRM users merge records constantly. When they do, the other system still points at an id that no longer exists, and the next sync recreates the record you just merged away. Handle tombstones and surviving ids explicitly.

Writes that succeed and change nothing

The API returns 200. The record is unchanged, because a validation rule rejected the field, or field-level permissions on the integration user exclude it, or a workflow overwrote your value a second later. Verify after write on a sample, and reconcile nightly, or you will trust a green log that describes nothing.

Personal data spreading faster than the policy

Each new destination is another place a deletion request has to reach. Mark personal fields in the manifest, keep the list of destinations current, and make deletion a first-class sync operation rather than a manual sweep somebody performs under time pressure. More detail in personal data in AI pipelines.

Zero errors is not the same as working

The most common way a CRM integration fails is by stopping. No exceptions are thrown, no alerts fire, the dashboard is green, and the data quietly diverges for three weeks. Alert on the absence of expected events and reconcile on a schedule, because a healthy silent sync and a dead one look identical from the outside.

What it costs to run once live

Three real cost lines, and a fourth that is usually zero. Most CRM integration work is deterministic, so there may be no model in the running system at all.

Cost lineWhat drives itHow it scales
InfrastructureQueue, workers, a small database for the conflict log and watermarks.With event volume. Modest for most companies until you pass millions of records.
CRM API tierCall allowance, which varies by product and tier. Heavy sync can push you to a higher plan.With records changed per day multiplied by calls per record.
MaintenanceSchema changes, API version sunsets, new fields, new systems joining the sync.With the number of connected systems and how freely people edit the CRM schema.
Model tokensOnly if a model handles fuzzy entity resolution or free-text normalisation. Often zero.With the volume of records that deterministic rules could not match.
If a vendor's quote has a large model cost on a CRM sync, ask exactly which decisions the model is making.

The number worth estimating before you commit is API calls, because that is the ceiling you hit first. Count calls per record honestly: a lookup, a dedupe check and a write is three, and any enrichment adds more. Put your own allowance in the box below, taken from your CRM's current limits documentation for your tier rather than from an article.

Daily API call estimate against your own allowance

Enter your CRM's documented daily allowance for your tier. The peak figure assumes a fifth of the day's volume lands in the busiest hour, which is normal for business-hours systems.

0Calls per day
0Calls per minute at peak
0Percent of your allowance used

Run it with backfill volumes too. A one-off load of half a million historical records at three calls each is one and a half million calls, which is why the backfill gets its own path and its own window. The connective layer underneath is shared with systems integration, so if you are building several of these it is worth building the queue once.

The cheapest optimisation is not syncing the field

Before adding a field, ask which system reads it and what decision changes because of it. Roughly a third of the fields on a typical wish list have no reader. Cutting them removes calls, conflicts, maintenance and personal data exposure at the same time.

How to tell whether you need this

The signal is not the number of tools. It is whether the same fact exists in two systems with no stated rule about which one is right, and whether anybody has stopped trusting the CRM because of it.

Signals. Four or more and the case is usually clear
0 of 6 done

If you ticked one or two, the fix is probably a written ownership map and an afternoon of configuration, not a build. Write the map anyway. It costs nothing and it is the deliverable that makes the eventual build cheap. Teams whose next step is pipeline hygiene rather than sync should look at sales pipeline automation instead.

Run the audit even if you buy nothing

Profile fill rates, duplicate rate on your natural key and last-writer per field. It takes a day with an export and some SQL, and most teams find one field being written by three systems.

How to start

Read access to both systems and one week is enough to tell you whether this is a mapping problem, a duplicate problem or an adoption problem. The three have very different price tags and only one of them is an integration.

  1. Send the list of systems, the objects that exist in more than one, and the fields people currently re-key by hand.
  2. We profile the data, propose the natural keys, measure the duplicate rate and draft the custody manifest for the contested fields.
  3. You get a scoped build with stages and a cost, or a recommendation to configure the connector you already own and stop there.
  4. If it proceeds, one low-risk field goes end to end first, so the architecture is proved on something nobody would miss.

The mechanics that keep this reliable overnight are covered in idempotency in automation, which is worth reading before you evaluate any vendor in this category.

Cite this

ChatGPTalker. CRM Integration: Sync, Field Custody and Conflict Rules. chatgptalker.com/services/crm-integration/

Questions we get asked

Is a native connector enough, or do we need a custom integration?
Start with the connector. It is cheaper, faster and maintained by somebody else. Move to a custom build when you hit a specific wall: a field it cannot map, a direction it cannot express per field, custom objects it does not support, or a conflict it resolves in a way that loses data. Keep a list of those walls, because that list is the specification.
Can you sync two CRMs, or migrate from one to another?
Both, and they are different jobs. A migration is a one-off with a cutover date, a reconciliation and a rollback plan. A two-CRM sync is permanent and needs a custody decision for every shared field, which usually surfaces the question of why two CRMs exist. That question is worth asking before commissioning the sync.
How do you stop duplicate records being created?
By deciding the natural key first and profiling its duplicate rate before any code is written. Matching is deterministic wherever possible: normalised email, cleaned domain, known exceptions listed. Fuzzy matching runs only above a confidence threshold and proposes rather than merges, because automatically merging two real customers is one of the few errors here that is genuinely hard to undo.
Does a CRM integration need AI at all?
Usually not, and we will tell you so. Sync, custody, dedupe and reconciliation are deterministic problems that deserve deterministic solutions you can test. A model earns its place in narrow spots: matching messy company names, classifying free-text notes, or summarising a support history onto a record. Everything else is code, and code is cheaper to debug at midnight.
What happens when the CRM schema changes?
The sync should fail loudly rather than guess. Unknown picklist values are rejected and logged instead of coerced, and a missing field raises rather than silently skips. Beyond that, the manifest is compared against the live schema on a schedule, so a field somebody added on Tuesday appears in a report rather than as a mystery three weeks later.
How do you handle deletions and data subject requests?
Deletion is a first-class sync operation. Personal fields are marked in the manifest, every destination is known because the manifest lists the readers, and a deletion propagates as a tombstone rather than a silent no-op. A request can then be honoured across all connected systems and evidenced afterwards, which is the part auditors ask about.

Tell us what is eating the hours.

Send the process, the volume and the tools it touches. You get a scoped plan with a build shape and a timeline, not a brochure.

Start a project