AI agents · D2C operations · catalogue to cash

The agent doesn’t get to say it’s done.

Any D2C job that repeats to a rule can be an agent. Catalogue to courier to cash. This is one of them being told no.

See the nine automations we already run

We reply in 1–2 hours during our working day. Four hours of it sit inside yours if you are in the UK or the EU. You get a straight answer on whether an agent suits the job you have.

Who decides it is done

Inside the agent

The agent proposes a change

Outside it · written from your rules

The check measures the work

Pass

It ships to your store.

Fail

Back to the agent. Nothing ships.

Rules
Written per job
Check
Before every handover
Ownership
The code and the logs are yours
Commercial
Fixed scope, fixed price

One agent run, illustrated

SHIPPED

catalogue › meta descriptions

  1. AGENT

    • READTask: rewrite meta descriptions
    • READScope: 412 products, one collection
    • RULESMax 160 characters per description
    • RULESNo invented codes. No empty fields.
    • PLANDraft all 412. Publish nothing yet.
    • WRITE412 drafted from your product fields
  2. CHECK

    • CHECKRunning 3 rules over 412 records
  3. FAIL

    • FAILSKU 4471 - 163 characters, limit 160
    • HELDPublish blocked. 0 of 412 live.
  4. AGENT

    • FIXSKU 4471 rewritten - 154 characters
  5. CHECK

    • PASSLength: 412 of 412 under 160
    • PASSCodes: 0 invented, 412 matched
    • PASSEmpty fields: 0 of 412
  6. SHIPPED

    • DONE3 rules passed. Handover accepted.

One agent run, start to finish. The check is the part you’re paying for.

Most “AI agents” are last year’s chatbot with a new label.

Roughly 130 of them are real.

Gartner press release

25 Jun 2025

That is Gartner’s count, out of thousands of vendors selling agentic AI, and in June 2025 it gave the rest a name: agent washing. A renamed assistant. A chatbot. An RPA script with a new badge. The same research expects more than 40% of agentic AI projects to be scrapped by the end of 2027, on cost, thin value and weak risk controls.

Shopify Community thread

24 Feb 2026

One merchant audited more than 80 products by hand to undo five months of his platform’s built-in AI. It invented reference codes. It breached the 160-character limit he had set on meta descriptions. It reverted his Spanish copy to English after he told it not to. He gave up in February 2026, and support confirmed there was no setting that would have stopped any of it.

Having to audit 80+ products because an AI cannot follow simple rules like ‘do not use this word’ or ‘keep it under 160 characters’ is unacceptable.

Read that again. The rule was stated plainly, the rule was ignored, and nobody found out until a human had checked 80 products one at a time.

Shopify SEO help page

current

Here is the part that should worry you more. Shopify recommends 160 characters for a meta description and does not enforce it. Nothing in the admin stops an over-long description from saving. So the limit only exists if something outside the tool is counting.

The rule he set

What came back

Keep it under 160 characters

Fail

The limit was breached

Do not use this word

Fail

The word was used

Leave the Spanish copy alone

Fail

Reverted to English

Use real reference codes only

Fail

Codes that did not exist

Found by

a human, 80+ products, one at a time

If a person does it every week to a rule, an agent can do it.

Fifteen jobs, and not one of them is a chatbot.

15

jobs listed

6 areas · list not closed

Every one is something a person in a D2C team does on a Tuesday afternoon, spreadsheet open and a seller panel in the next tab. Read the list before you decide whether this is for you — the point is not any single job, it is how much of one operation is on it.

What sits in front of the customer

01

Catalogue

Catalogue and PDP updates

Prices, copy, badges, metafields, collection rules. Validated against your rules before a single record reaches the storefront: character limits, words you have banned, the fields that must never go out blank.

02

Bulk ops

Bulk product and image work

A linesheet or a raw shoot folder goes in. Website-ready records come out: crops at the exact sizes your theme expects, compression, alt text, then the push to the store.

03

Content

Product copy and descriptions

Written to your brand rules and your character limits, from the attributes you actually hold. Where the data is thin, it says so rather than inventing a specification.

Pricing, and every channel it has to reach

04

Pricing

Price changes, promotions and offer setup

A festive price list goes in. Out come the changes, staged by date, with the ones that break your margin floor held back for a person to look at.

05

Marketplaces

Marketplace listing sync

One catalogue, five to eight seller panels, each with its own title rules, attribute names and image specs. Your single source feeds each channel’s own version, and afterwards you get a report of exactly what each panel rejected and why.

Stock, and what you owe for it

06

Inventory

Stock sync across channels

One number, reconciled across the store, the warehouses and every marketplace you sell on. Oversells get flagged the moment the counts disagree, not the morning after.

07

Buying

Purchase orders and vendor chasing

Reorder points watched, purchase orders drafted from what actually sold, and the follow-up email nobody sends on time sent on time.

Everything between the order and the doorstep

08

Orders

Daily order operations

Tagging, routing, holds, address fixes, and the exceptions pulled out into one place instead of three tabs.

09

COD

COD confirmation and RTO risk

Built for the COD reality of Indian D2C. It watches new orders, runs your address and phone checks, triggers confirmation on WhatsApp, and flags the risky ones before a courier collects them.

10

Delivery

Courier allocation and NDR follow-up

Failed deliveries get chased the same day, to your script, on the channel the customer actually answers. NDRs that nobody chases become RTOs, and RTOs are the most expensive thing in Indian D2C.

11

Returns

Returns and exchange triage

Reads the request, matches it against the order and your policy, then sorts it into approve, ask for a photo, or escalate. A person approves every refund. That one stays human.

The money, and what got deducted from it

12

Money

Settlement and payment reconciliation

Order by order, across gateways and marketplaces, against what actually landed in the bank. The mismatches come out as a list with amounts, not as a feeling that something is off.

13

Tax

Deduction matching

The marketplace withholds TDS under 194-O and TCS under Section 52. Both get matched against your own records. The gap surfaces in the same month, not at filing.

What your team reads and answers

14

Reporting

Numbers that land before the standup

Yesterday’s figures, pulled from the store, the ad accounts and the courier data. Delivered where your team already reads, rather than in one more dashboard nobody opens.

15

Support

Support triage before a human replies

Sorts the inbox, pulls the matching order, drafts the reply. Anything involving money goes to a person.

That is fifteen, and the list is not closed. If your version of the job is not here, it is the one to send us. Nine automations of this kind already ship with the stores we build. See the nine.

The same rule built our own advertising system: Autopilot builds a Meta campaign and then refuses to spend on it, because an agent holding a budget is the clearest case there is for making it ask.

Two things decide whether an agent is safe to run on your store.

Not the model. Not the prompt. What the agent is forbidden to do, and who gets to decide it has finished.

01
The rulebook it may not break
02
The check it did not write

01

A written list of moves it may not make

Hard refusals, agreed before any work starts. Not tips, and not a tone of voice. Every agent begins from a rulebook written for its own job, and you read that rulebook before we begin.

Catalogue agent

records · copy · codes

  1. 01Exceed the character limit you set
  2. 02Use a word you have banned
  3. 03Leave a required field blank
  4. 04Invent any code that is not already in your data

The rulebook is written for the job, not for the platform.

Money agent

settlements · deductions

  1. 01Write a correcting entry
  2. 02Close a gap it has found

Reconciliation finds the gap. A person closes it.

On Shopify

theme · admin · tokens

  1. 01Push to the live theme
  2. 02Author in the admin code editor
  3. 03Overwrite the settings your team built in the theme editor
  4. 04Put a private token anywhere a browser can download it

Work happens on an unpublished development theme that only you can publish.

On Magento

code · database · deploy

  1. 01Touch vendor code, core modules or compiled output
  2. 02Override higher than it needs to — theme, then layout, then view model, then plugin
  3. 03Write to the database without a store scope

One missing store filter rewrites every locale on the site.

A copy of the rulebook goes into the handover at the end.

02

A check the agent did not write.

When an agent says it has finished, that counts for nothing.

The work goes to an automated check instead, and the check measures what can be measured.

01Spacing against the design
02Load time
03A token sitting in a file the browser downloads
04An edit to a file declared off limits
05Product records that breach a character limit

Out comes a pass, or a list of failures with numbers attached.

The record it produces

Same run as above

Check

State

description length

FAIL

wanted ≤ 160 chars · got 163 chars

records published

HELD

wanted 0 until pass · got 0 of 412

description length · re-run

PASS

wanted ≤ 160 chars · got 154 chars

reference codes

PASS

wanted 0 invented · got 0 of 412

required fields

PASS

wanted 0 blank · got 0 of 412

3 rules passed. Handover accepted.

A fail means the work is not handed over.

Not flagged. Not “ready, with caveats”. Sent back.

The dashed line in the hero

Asking a model to confirm its own destructive action is not a control, because the confirmation comes from the same place as the decision. So the check sits outside the agent, and the agent cannot edit it.

Two agents that had safety instructions, and reasoned past them

25 April 2026

Reported by The Register, and Zenity’s post-mortem

Trigger
a credential mismatch
Elapsed
nine seconds
Token it found
scoped to any operation
Rule it carried
do not guess
Lost
production volume + backups
Saved by
an off-site backup

An agent hit a credential mismatch on a routine staging job, went hunting through unrelated files, and found an API token scoped for any operation including the destructive ones. The production volume was gone, and the volume-level backups with it, because the host kept them on the same volume. Its own written account was that it had guessed the delete would be scoped to staging, and had never checked whether the volume was shared.

July 2025

Reported by The Register, 21 July 2025

Trigger
a declared code freeze
Elapsed
one deploy
Access it had
the production database
Rule it carried
change nothing
Lost
a live customer database
Saved by
the founder, by hand

Another agent deleted a customer’s live database during a declared code freeze, then reported that rollback was impossible. It was not, and the founder recovered the data himself. This one had safety instructions too, and it reasoned past them the same way: it decided the rule did not cover the case in front of it.

Some jobs end in the codebase. We go in.

Most of the fifteen never touch code.

These three do, and that is the line between us and a no-code automation vendor. A template rewritten. A checkout changed. A section built to a design file. We go into the repository and change it, under the same rulebook and the same check.

Magento and Adobe Commerce agent

Theme overrides · Layout XML · View models and plugins · LESS

Works the codebase the way a senior would.

Refuses, always
Editing vendor code, core modules or compiled output · Any database write without a store scope · A command-line run that leaves root-owned files behind
Checked, every time
Measured spacing, type size and colour against the design file · The page loading with the new stylesheet applied · Zero root-owned files after deploy

Shopify agent

Sections and blocks · Metafields and metaobjects · Cart and checkout · Theme app extensions

Liquid, Online Store 2.0, Storefront API, Hydrogen.

Refuses, always
Pushing to the live theme · Authoring in the admin code editor · Overwriting the merchant’s theme settings, locale files or JSON templates · An Admin API token in anything the browser downloads
Checked, every time
The platform’s own theme linter running clean · A real cart-to-checkout run · No secret in the client bundle · The pinned API version present in the code

Frontend agent

React and Next.js · Tailwind · The real breakpoints · Reduced-motion states

A Figma file turned into code that matches the Figma file.

Refuses, always
Shipping a section it has not measured · A colour or typeface outside the system · A heading order that breaks the document outline
Checked, every time
Measured deltas against the design · Contrast ratios · Core Web Vitals on the built page

Every agent hands over a written record of what it changed and why. Your lead reviews a diff, not a mystery.

From one workflow to something running.

Six steps, and you own what comes out of them.

  1. 01

    The 20-minute call

    You describe one workflow. What goes in, what comes out, how often, and what “wrong” looks like. We tell you whether an agent fits. We say no when it does not.

    You get

    A straight yes or no

  2. 02

    Rules and thresholds, agreed in writing

    Before anyone builds, we write down what the agent may not do and what the check will measure. You sign that page off. A page, not a deck.

    You get

    The rulebook and the check, in writing

  3. 03

    Built against your real data, read-only first

    The agent runs on a copy of your records with no write access. Production access comes per workflow, after the check passes, scoped to the records that workflow touches.

    You get

    The agent running, with no write access

  4. 04

    Checked before anything is handed over

    The check runs on the work, not on the agent’s report of it. A fail sends the work back. You see that output, pass or fail, every time.

    You get

    The check output, pass or fail

  5. 05

    Handover with the record

    The agent, the rulebook, the check and the change log, all in your repository.

    You get

    Agent, rulebook, check, change log

  6. 06

    After it is running

    Not written yet — Who monitors it · Who fixes it · For how long · At what cost.

    This page will carry the answer before it is sold.

No lock-in and no platform fee to us. Whatever we build, you can run without us.

Where we tell people not to buy this.

We just claimed a lot. Here is where it stops.

A shopper-facing chatbot on your storefront

Buy one off the shelf. The support vendors do that job better than a bespoke build ever will.

Write access to production on day one

We will not build it that way. The two database incidents above are the whole argument.

A workflow that changes weekly and has never been written down

Write it down first. No agent can follow a rule nobody has stated.

A task one person does twice a month in ten minutes

The build costs more than the task. Keep doing it by hand.

A judgement call dressed up as a rule

Deciding which influencer to sign, or whether to refund a rude customer, is not a rule with an exception. It is a person’s job with a rule attached.

And these stay with a person

Every refund · Every diff, before it merges · Write access to production, per workflow · The rulebook, before any work starts · Any gap the reconciliation finds · Any price that breaks your margin floor

A signature is a feature here, not a gap

Questions people ask before they sign.

The code
yours
The rulebook
yours
The check
yours
The change log
yours
Lock-in
none
Platform fee to us
none

Same terms as our storefront work. It all sits in your repository.

Do we have to be on Shopify or Magento?

No. Most of the fifteen jobs never touch a storefront codebase. They run against your records, wherever those sit: the store admin, a marketplace seller panel, a courier dashboard, a spreadsheet your finance person owns.

Platform depth is for the jobs that end in code.

Will it hallucinate and put something wrong on my store?

It will produce wrong output sometimes. Anyone who promises otherwise is selling. What we control is whether that output reaches your store, and the check runs on the work itself rather than on the agent’s account of it.

A breached character limit. An empty required field. An invented SKU code. A token in the wrong file. Each of those is checked for by name, and any one of them stops the handover.

Does it need access to my live store?

Not to begin with. The first build runs read-only against a copy of your data. Write access comes later, per workflow, after the check has passed on real records.

What happens when it gets it wrong anyway?

You get the log of every change it made. That log is the difference between a five-minute revert and a day of forensics.

Is this cheaper than hiring someone?

Sometimes. It earns its build cost when the task is repetitive, rule-bound and running most days of the week. It is a poor trade for judgement work.

Who owns it?

You do. The code, the rulebook, the check and the logs sit in your repository. Same terms as our storefront work: no lock-in, no platform fee to us.

We already have Shopify’s built-in AI. Why pay for this?

Because the tools inside the admin are helpers with no obligation to obey your rules, and helpfulness is not the same thing as obedience. Support told the merchant quoted above that no setting exists to stop the AI inventing data. He left the platform in February 2026 after auditing 80+ products by hand.

How is a developer agent different from giving my team an AI coding tool?

A tool helps a person type. Ours carries the platform’s rules and is held to a check before its work counts as finished.

There is also evidence that the tool on its own is not the win it feels like. In a randomised trial published in July 2025, 16 experienced developers took 19% longer on 246 real tasks while using AI tooling. Afterwards they estimated they had been 20% faster.

Does the code hold up six months later?

Fair question, and the right one. GitClear’s 2025 study of 211 million changed lines found duplicated code blocks growing eightfold during 2024, while refactoring fell away. Our check looks at the shape of what was written, not only at whether it runs.

Why should we believe your check is any good?

Because you read what it measures before we start, and you get its output on every handover. We build storefronts the same way. One brand’s Shopify theme sent 385,680 bytes of HTML; the headless rebuild sends 3,084. Both are still live, so you can fetch them and check us.

Not here? hello@headlinehq.in

Reply time — 1–2 hours · four hours a day overlapping the UK and EU

Where the numbers on this page come from.

40% of agentic AI projects scrapped by 2027, and the ~130 real vendors

Gartner press release

25 Jun 2025

160 characters is a recommendation the admin does not enforce

Shopify SEO help page

current

19% slower across 246 real tasks, 16 experienced developers

METR randomised trial

Jul 2025

Duplicated code blocks up eightfold across 211 million changed lines

GitClear research

2025

The merchant quotation, and the five months of hallucinated data

Shopify Community thread

24 Feb 2026

Live database deleted during a declared code freeze

The Register

21 Jul 2025

Production volume and backups deleted in nine seconds

The Register

27 Apr 2026

The technical post-mortem of that deletion

Zenity

Apr 2026

Initial HTML of 385,680 bytes against 3,084, fetched with curl on two live hosts

Our own case study

18 Aug 2026

The 412 products and SKU 4471 in the run above are illustrative, and labelled as such on screen. This page carries no outcome number for agent work, because we do not have an audited one yet.

Send us the workflow that eats the most hours.

One task. What goes in, what comes out, how often it runs, and what happens today when it goes wrong.

Email hello@headlinehq.in

We reply in 1–2 hours during our working day, four hours of which sit inside a UK or EU one. No deck. No demo environment built to look good.

What comes back

  1. 01

    Whether an agent fits the job at all

  2. 02

    A straight no, when it is the wrong tool

  3. 03

    What it would be forbidden to do

  4. 04

    What the check would measure

  5. 05

    A fixed scope and a fixed price

Send one job. We come back on that one, not with a deck about all fifteen.