# Language alternates — full dumps in other locales: # RU: https://inite.ai/llms-full-ru.txt # PT: https://inite.ai/llms-full-pt.txt # ES: https://inite.ai/llms-full-es.txt # INITE — Full Content Dump > Turn chaos into profit with AI automation INITE builds shared infrastructure for intelligent operations and applies it through INITE Solutions and vertical products. We deploy 1-3 workflows in production in 2-4 weeks and hand them to the client's own team; if diagnostics cannot show a payback, we do not build. Source URL: https://inite.ai Locale: en Generated: 2026-09-09T05:39:45.334Z --- # Choosing Between Two Automation Vendors Who Both Said Yes URL: https://inite.ai/en/blog/choosing-between-two-automation-vendors Date: 2026-09-08 Author: Anton Fenix Category: Operations Tags: Operations, Procurement, Comparison, ROI ## Direct Answer When two proposals both look fine, price is the worst way to choose, because the cheaper number usually describes less work rather than the same work done cheaper. Five questions separate them. What baseline did you measure, and how. What is the total across the build plus twelve months of running it. What would you refuse to build for us. Who owns the code and the account keys on the day we part. And what happens when a system you integrate with changes shape. Every one of those has a right answer that costs the vendor something to give, which is exactly why it tells you more than the price does. ## Key Facts - A cheaper quote usually describes less work rather than the same work at a lower price. - Build price plus twelve months of running cost is the only total two proposals can be compared on. - A vendor who cannot name something they would refuse to build has not thought about your process yet. - Ownership of code and credentials is decided on the day of the handover or on the day of the argument. ## The problem with two reasonable numbers Two vendors, two proposals, both credible, one cheaper. The instinct is to treat the difference as efficiency. It almost never is. Scope is invisible in a total. One proposal priced the exception path, the reconciliation, and a month of tuning after launch. The other priced the path where everything goes right, and will meet the rest as change requests. Neither is lying. They are describing different jobs, and the totals hide which is which. ## Five questions, none about price **What baseline did you measure, and how?** A vendor who cannot say where the current number came from will be free to invent the comparison later, when the project needs a result. The before figure has to exist before the after does, which is the whole argument in [what a process audit must produce](/en/blog/process-audit-before-automation). **What is the build plus twelve months of running?** Usage at your real volume, hosting, monitoring, and the hours a person spends each month on whatever the workflow hands back. That last item is missing from most quotes and is often the biggest. One sum, and most of these decisions resolve themselves - the same arithmetic that [reading a quote](/en/blog/how-to-read-an-automation-quote) is built on. **What would you refuse to build for us?** The useful answer is specific: a decision that should stay with a person, a data source too unreliable to sit under a workflow, a step whose volume does not repay the work. A vendor who would build all of it has looked at your budget rather than at your process. **Who owns the code and the keys?** Where it lives, who can read it, whose accounts hold the credentials, and what you have on the Monday after you stop paying. Cheap to agree now, unpleasant to discover later. **What happens when an integration changes shape?** Every workflow depends on somebody else's release schedule. The good answer names how a silent change would be caught, which is a question [worth asking before the build](/en/blog/when-the-integration-changes-underneath-you) rather than after the first bad month. ## What the answers are actually measuring All five cost the vendor something to answer honestly. That is the point. A price is free to state and free to be wrong about; a refusal, an ownership clause and a running figure in writing are commitments, and [commitments sort vendors in a way totals cannot](/en/answers/how-to-choose-ai-automation-agency). ## When both answer well Then you have a real choice rather than a trap, and the remaining differences - sequence, who is in the room, how change is priced - are ones you are qualified to judge. Choose the one whose refusal list you agree with. That list is the closest thing available to a preview of how they will behave when something goes wrong, which it will. ## FAQ ### Why is the cheaper proposal usually not the cheaper job? Because scope is invisible in a total. One vendor priced the exception path, the reconciliation step and a month of tuning after go-live; the other priced the happy path and will treat everything else as a change request. Both numbers are honest descriptions of what their author intends to do, and the difference between them is not efficiency but coverage. The way to see it is to compare the two proposals against a brief you wrote yourself, so what is missing from one of them is visible as an absence rather than as a saving. ### What does the running cost actually include? Model or platform usage at your real volume rather than at a demo volume, the hosting, the monitoring, and the hours somebody spends each month on the exceptions the workflow hands back. That last item is the one left out of most quotes and is often the largest. Ask for a monthly figure in writing, multiply by twelve, and add it to the build price. That single sum makes most vendor decisions obvious in one direction or the other, and it is cheap to obtain. ### Why ask what a vendor would refuse to build? Because the answer is evidence about whether they have looked at your process or only at your budget. A vendor who has done this before will name something specific - a decision that should stay with a person, a data source too unreliable to build on, a step whose volume does not justify the work - and will explain why. A vendor who says they can do all of it is describing their sales position rather than your operation, and the parts they should have refused will surface during acceptance instead, when they are expensive. ### What ownership questions matter on day one? Where the code lives and who can read it, whose accounts hold the API keys, who can revoke access, and what you receive if the relationship ends next quarter. These are cheap to agree before a contract and unpleasant to discover afterwards, when the answer is whatever the vendor's default happens to be. The test is simple: ask what you would have on the Monday after you stop paying, and expect a specific answer rather than a reassurance. --- # What the Break Map Shows, and What It Decides URL: https://inite.ai/en/blog/protocol-break-what-the-map-shows Date: 2026-09-03 Author: Mikhail Savchenko Category: Methodology Tags: INITE Protocol, Process Audit, Methodology, ROI ## Direct Answer The Break stage produces three things a client keeps: a map of how the work actually moves, a cost-of-chaos figure attached to the candidate workflows, and a matrix that scores each candidate on feasibility against return. The map is the input. The other two are the decision. Without them, the choice of what to automate first is made on whoever complained loudest, and the before number gets reconstructed after the build from memory, which always flatters the result. Break runs in weeks 1-2 and is allowed to end the engagement. ## Key Facts - Break runs in weeks 1-2 of the engagement and is allowed to end it, with the diagnostic deposit refunded and the reasons written up. - The cost-of-chaos figure is built from loaded hourly cost and volumes pulled out of the systems, not from headline salaries or from a conversation. - Every candidate workflow is scored twice, on feasibility and on return, and the second score is what orders the build queue. - The client keeps the map, the baseline and the matrix whether or not a single line of code is written afterwards. ## The deliverable people expect Ask most operators what an automation audit produces and they will describe a diagram. Boxes, arrows, swimlanes, a slide deck built around it. That is the input, not the output. A map tells you what the process is. It does not tell you what to do, and the gap between those two is where every bad automation decision gets made. ## Two numbers turn a map into a decision The first is the cost of chaos: what these specific workflows lose per week to rework, waiting and dropped handoffs. Loaded hourly cost, volumes from the systems, a window long enough to catch a bad month. The second is a score per candidate, on two axes. How feasible is this to automate reliably, and how much of that weekly cost would it actually recover at conservative assumptions. Neither is sophisticated. Both are absent from most proposals, and the absence is not an oversight. A proposal without a measured baseline cannot be checked afterwards, which is convenient for the party writing it. ## What goes wrong when the stage is skipped Three things, in a reliable order. **The wrong process gets automated first.** Without a return score, selection defaults to whoever complained most recently or whose process is easiest to describe in a meeting. Those correlate weakly with cost. **The before number gets invented afterwards.** This is the expensive one. Nobody measured the old cycle time, so when the new one is reported at four hours, the old one becomes "about a day" - a figure produced by the same people who need the project to have worked. The [audit before automation](/en/blog/process-audit-before-automation) exists to make that reconstruction unnecessary. **The engagement cannot end in no.** A diagnostic with no arithmetic in it has no mechanism for concluding that the build is not worth doing, so it never concludes that. Every audit returns yes, and the yes means nothing. ## The map still has to be honest The measurement only works if the map underneath it describes the real route, workarounds included. The shared spreadsheet nobody officially uses is load-bearing, and it is usually where the contradictions that break automation are hiding. How that map gets drawn is a subject of its own, covered in [why most process maps are useless](/en/blog/protocol-break-mapping-the-process). ## What the client keeps The map, the baseline and the matrix, in writing, regardless of what happens next. That is the design: the stage has to be worth its own fee even when it kills the project, or the incentive to find a project quietly returns. Break is one of [six stages](/en/protocol), and the ones after it depend on this output being real. The full sequence, walked through a single deployment, is in [the protocol applied end to end](/en/blog/inite-protocol-6-stages-applied). ## FAQ ### Why is the map not the deliverable? Because a map answers the wrong question. It tells you what the process is, and the question on the table is what to do about it, which needs two more artefacts. The first is a number: how much this process costs per week in rework, waiting and missed handoffs, calculated across the candidate workflows rather than across the company. The second is a ranking: which candidate returns the most for the least build risk. A vendor who hands you a beautiful diagram and a proposal has skipped the step where the diagram turns into an argument, and you are being asked to approve a build on aesthetics. The map is the evidence. The cost figure and the matrix are the finding. ### What makes the cost-of-chaos number trustworthy? Its construction, and the fact that it is written down before anyone knows what it will justify. It is built from loaded hourly cost rather than headline salary, because the hour a person spends on rework costs the business what that hour costs to provide, not what appears on the payslip. Volumes come out of the operational systems for a period long enough to include a bad month, not out of an estimate offered in a meeting. And it is scoped to the candidate workflows only, which keeps it small and specific and stops it drifting into the kind of company-wide inefficiency figure that cannot be checked against anything. The test of the number is simple: it is the baseline every later return claim gets measured against, so it has to be one you would be content to be held to. ### What does the priority matrix actually score? Two axes, kept deliberately crude. Feasibility asks whether the inputs are structured enough, the systems reachable enough, and the rules stable enough that software can do this reliably rather than mostly. Return asks how much of the measured weekly cost this candidate would actually recover, at conservative assumptions rather than optimistic ones. Scoring both matters because the highest-return candidate is frequently the least feasible one, and a plan that ignores that ships a project which is exciting for three weeks and abandoned in the fourth. The output is a queue, and the point of the queue is that the first item is defensible to someone who was not in the room. ### How often does Break end the engagement? Often enough that the gate is real, which is the only property that matters. If the arithmetic at conservative inputs does not clear positive, the diagnostic deposit is refunded, the reasons are written up, and the engagement stops there. This sounds like a marketing line and it is the opposite: a filter that never rejects anything is not a filter, it is a formality, and every consultancy that has one which never fires is running the second thing while describing the first. The client still leaves holding the map, the baseline and the matrix, which are worth having on their own and which make the next automation attempt, whoever runs it, materially cheaper. --- # The Automation Brief an Operator Should Write First URL: https://inite.ai/en/blog/automation-brief-template-for-operators Date: 2026-09-01 Author: Mikhail Savchenko Category: Operations Tags: Operations, Procurement, Process Audit, Automation ## Direct Answer Write the brief before you speak to anybody. Seven items: the process in one sentence, the volume per week from a system rather than from memory, who touches it and for how long, what finished means, what must never be automated, what it costs you now, and what evidence you will accept that it worked. An afternoon of work, and it changes what comes back. Vendors given a brief quote the same thing and can be compared; vendors given a conversation each quote a different thing, and the cheapest number wins by describing the least. The brief is also the only artefact that survives if you decide not to build. ## Key Facts - Three proposals written from three different conversations cannot be compared, because they are not quoting the same work. - Volume taken from a system rather than from memory is the single item that changes a quote most. - The line that says what must never be automated is the one vendors most often leave out and clients most often mean. - The brief keeps its value if the decision is not to build, which is more than most discovery documents manage. ## Why the brief comes first Ask three vendors to quote for automating your order intake and you will get three prices for three different jobs. Not because anyone is being dishonest, but because each of them heard a different half-hour of description and filled the gaps with what they usually build. The brief closes the gaps on your side of the table. It costs an afternoon, and it [turns a set of incomparable proposals into a set of comparable ones](/en/answers/how-to-choose-ai-automation-agency). ## The seven lines **One. The process, in one sentence.** From what event to what outcome. "A rental inquiry arrives by phone, WhatsApp or the web form, and ends when the equipment is on a van with a signed agreement." If it takes three sentences, you have two processes. **Two. Volume per week, from a system.** Not a typical day - the count for a period long enough to contain a bad week. This is the line that moves a quote most, and the one most often supplied from memory. [What to automate first](/en/blog/what-to-automate-first-in-a-small-company) is decided by this number more than by anything else in the document. **Three. Who touches it, and for how long.** Roles, not names, and minutes per item rather than a share of a day. Two people at twenty minutes is a different problem from six people at four minutes. **Four. What finished means.** The state the process is in when it is done, written so a stranger could check it. This is the acceptance test, and writing it now prevents the version where finished means whatever the delivered system happens to do. **Five. What must never be automated.** Refunds above a threshold, anything clinical, anything that goes to a regulator, anything a customer would experience as a machine deciding their case. This is the line vendors leave out and clients turn out to have meant all along. **Six. What it costs now.** Hours per week at a loaded hourly cost, plus the rework you can name. Not the whole company - this process. The arithmetic for producing a number you can defend is in [the four questions that break most ROI claims](/en/blog/roi-math-for-automation-projects). **Seven. What evidence you will accept.** The measurement that will settle it in ninety days, and where the number will come from. Agreeing it now is what stops the after-the-fact comparison against a remembered baseline. ## What it changes on the other side A vendor given this document quotes the same job as every other vendor given it. That is the whole point. Where they still differ - approach, sequence, what they refuse - the differences are real ones you can judge, and [reading the quote itself](/en/blog/how-to-read-an-automation-quote) becomes a much shorter exercise. It also tells you something about the vendor. An experienced one will add to line five and argue with line two. One who treats the brief as scope being taken away is telling you where their margin lives. ## The version where you do not build Sometimes the sixth line comes out small. That is a result, not a wasted afternoon: it arrives before a build rather than after one, and the same page is still a written description of one process, its volume and its real cost. The next attempt starts from it, whoever runs that attempt and whenever it happens. ## FAQ ### Why write the brief before talking to vendors rather than after? Because a proposal written from a conversation describes the vendor's understanding of what you said, and three vendors heard three different things. That makes the quotes incomparable in a way that is invisible: they are all reasonable, all confident, and all pricing a slightly different job. Writing the brief first fixes the scope on your side of the table, so the differences that remain between proposals are differences in approach and price rather than in what is being built. It also moves the scoping decision away from the party whose income scales with the scope. ### How precise does the volume number need to be? Precise enough to have come from a system rather than a memory, and covering a window long enough to include a bad week. An operator asked how many orders they process will answer with a typical day, which is almost always the number they wish were typical. Pulling the count from whatever system already records it takes twenty minutes and routinely moves the answer by a third in one direction or the other. It is the item that changes a quote most, because volume decides whether a rules-based approach is enough or a model is needed at the front door. ### What belongs in the line about what must never be automated? The decisions where being wrong is expensive and cannot be undone quietly: a refund above a threshold, a clinical or legal judgment, anything that goes to a regulator, anything a customer would experience as a machine deciding their case. Writing them down before the build is what turns them into a specification rather than an argument during acceptance. It is also the fastest way to tell whether a vendor has done this before, because an experienced one will add to your list and an inexperienced one will treat it as scope being taken away. ### What if the brief shows the process is not worth automating? That is the brief working. The measured cost per week is the same number every ROI claim will later be checked against, and when it is small the honest conclusion is available immediately rather than after a build. The document does not lose its value in that case: it is a written description of one process, its volume and its true cost, which makes the next attempt - by anybody, in any year - substantially cheaper than the first one was. --- # When the Integration Changes Underneath You URL: https://inite.ai/en/blog/when-the-integration-changes-underneath-you Date: 2026-08-28 Author: Mikhail Savchenko Category: Operations Tags: Operations, Automation, Workflow, Process Audit ## Direct Answer Most automations that stop being trusted did not stop working. Something upstream changed shape - a field renamed, a payload that gained a level of nesting, a status list that gained a value nobody mentioned - and the workflow carried on producing output that looks defensible and is wrong. The loud failure is the cheap one: it stops, somebody notices, it is fixed the same day. The expensive one runs for six weeks before a person compares two numbers by hand. Three cheap defenses catch it: one record pushed end to end every morning, assertions on the shape of what arrives, and an alert on the distribution of results rather than on exceptions. ## Key Facts - The failure that costs money is the one that does not raise an error, because nothing tells anybody it happened. - A canary record pushed through the whole path once a day is the cheapest detector of a silent change. - Checking the shape of an incoming payload catches renames and new nesting; checking only the status code does not. - An alert on the distribution of outcomes notices a workflow whose results quietly shifted, which an exception alert never will. ## The failure that does not announce itself An automation that stops is a small problem. It stops, somebody notices within the hour, the cause is obvious because the last thing that changed is the thing that broke. The expensive failure is the one that keeps running. A field is renamed upstream, the code that reads it gets nothing, and nothing is a legal value in most processes - an empty note, an unset flag, a missing second line of an address. Nothing raises. The workflow produces records that look exactly like last week's records, and it does that for as long as it takes somebody to compare two numbers by hand. ## What actually changes Four shapes account for most of it. A field is renamed or moved, usually as part of a tidy-up nobody thought was external. A payload gains a level of nesting when a vendor adds a wrapper for pagination or metadata. An enum gains a value - a new order status, a new document type - and the branch that handles the known values silently drops the unknown one. An API version sunsets, and the fallback turns out to be an older shape rather than an error. None of these are outages. Every one of them is a Tuesday afternoon in somebody else's release notes. ## Three defenses, all cheap **One record, end to end, every morning.** A synthetic item that goes the whole way and is checked at the far end against a known answer. It exercises the joins between systems, and the joins are where drift lives. Four separately healthy services can still be handing each other something that changed. **Assert on shape, not on status.** A two hundred with a parseable body is not evidence that the body means what it meant last month. Check that the fields you read are present and typed as expected, and fail loudly when they are not. This is ten lines and it converts a silent wrong answer into a visible stop. **Alert on distribution, not on exceptions.** If one in twenty items took the manual path last month and one in six takes it today, that is the signal - and no exception was raised to produce it. This only works if the ordinary numbers were written down first, which is the argument [a measured baseline](/en/blog/roi-math-for-automation-projects) makes for itself. ## Why this is a scoping question, not a maintenance one Every integration is a dependency on somebody else's release schedule. That does not make it a bad idea; the rental delivery in [where the four hours go](/en/blog/order-processing-equipment-rental) reads availability from a system we do not control, and it still pays for itself. It makes the dependency a thing to price. The practical version: when a workflow is scoped, list what it reads from outside itself, and for each one say what happens if the shape changes. Most answers will be "it stops, and that is fine". The ones where the answer is "it keeps going and we would not know" are the ones that need a canary before they ship, not after the first bad month. ## The part nobody wants to name [A workflow that no person is responsible for](/en/pricing) is a workflow being checked by whoever happens to look, which in practice means after a customer complains. The name goes in the handover document with everything else, and it belongs to somebody who has opinions about the design rather than to whoever was free that week. The argument for treating that as a stage rather than an afternoon is in [what a process audit must actually produce](/en/blog/process-audit-before-automation), and the reason it cannot be added later is that the person who forgot what was confusing cannot write it down. ## FAQ ### Why does an integration change break things silently? Because most integrations are checked for success rather than for shape. The call returns two hundred, the body parses, and the workflow proceeds. If a field was renamed, the code that reads it now gets nothing and treats nothing as an ordinary empty value - a blank customer note, a missing second address line, an unset priority. Every one of those is a legal value somewhere in the process, so nothing raises. The workflow keeps producing records that look exactly like the ones it produced last week, and the only signal is that the outcomes have drifted, which nobody is watching for. ### What is a canary record and why does it work? One synthetic item pushed through the entire path every morning - created, routed, transformed, delivered - and checked at the far end against a result that is known in advance. It works because it exercises the joins between systems rather than any single system, and those joins are where drift lives. A monitoring check that pings each service separately will report four healthy services while the thing they hand to each other has changed underneath. The canary costs one record a day and answers a question none of the individual health checks can. ### How do you tell drift from ordinary variation? By having written down what ordinary looks like before the automation shipped. That is what the baseline stage exists for: volumes per week, the share of items that take the exception path, the spread of processing times. Drift shows up as a change in the shape of those numbers, not in their level - the mean holds and the tail moves, or the exception rate steps from one in twenty to one in six overnight. Without a baseline the same numbers are unreadable, because nobody can say whether this week is unusual. ### Whose job is it to notice? Somebody named, before the workflow ships. A process with no owner is checked by whoever happens to look, which in practice means it is checked after a customer complains. The name belongs in the handover document alongside what the workflow does, what it must never do, and which alert means what. That is why the handover stage is a stage rather than an afternoon at the end of the build - a system nobody has been made responsible for is a system running on the assumption that nothing will change, and something always changes. --- # How to Read an Automation Quote Before You Sign It URL: https://inite.ai/en/blog/how-to-read-an-automation-quote Date: 2026-08-24 Author: Anton Fenix Category: Operations Tags: Operations, Procurement, Automation, Strategy ## Direct Answer Read the structure before the number. A quote worth signing names the process rather than the technology, states the volume it assumes, itemises integration per system, carries a monthly running figure alongside the build price, says what stays with a person, and names who owns the result after handover. What is missing is the informative part: a single-line total hides which assumptions can move, and a quote with no monthly figure is missing twelve numbers you will pay in the first year. The four running costs left out most often are model spend, the human handling exceptions, maintenance when surrounding systems change, and the process owner's continuing time. ## Key Facts - A quote with no monthly running figure is missing 12 numbers you will pay in the first year. - 4 running costs are left out most often: model spend, the human in the loop, maintenance, and the owner's time. - We deliver 1-3 workflows in 2-4 weeks, so a single-workflow quote measured in months is describing different work. - Payback typically lands at 3-6 months, and that cannot be checked at all without a monthly cost. ## Read the shape before the number Most quote reviews start at the total and work backwards, which is the wrong order. The total is a conclusion. The structure is the evidence. A quote is a description of scope wearing a price, and the first thing to check is whether the scope is defined by the process being automated or by a list of deliverables that could describe almost anything. "AI integration, discovery, implementation, testing, deployment" describes every project ever quoted. "Booking intake across four channels, availability resolved against asset state, agreement generated from your approved template, dispatch scheduled" describes one. ## What should be on the page | Line | Why it matters | | --- | --- | | Measurement or discovery, priced separately | A quote without one is a guess about an uncounted process | | Build, tied to a named process | Not to a technology, which could mean anything | | Integration, per system | Estimates go wrong here, and one combined figure hides which system is the risk | | Monthly running cost | The most commonly absent line, and the one that decides payback | | Handover, and what documentation exists | Decides whether you own the result or rent it | | Support terms, with a response time | Otherwise "we'll be around" is the whole commitment | Any of these missing is a question rather than a deal-breaker. All six present means the quote can be compared to another quote with all six, which is [the only way a price comparison means anything](/en/answers/ai-automation-cost). ## What a single-line total hides Which assumptions can move. Automation costs are dominated by volume and by integration surface, and at quoting time both are estimates. Itemised, you can ask what happens at half the assumed volume, or what the price becomes if the fourth system needs a different approach. The answers show you where the risk sits. Collapsed into one number, every later change becomes a renegotiation from a position where you cannot tell which part moved. Single-line quotes are more often lazy than dishonest. The effect on you is the same either way, and asking for a breakdown costs nothing and is refused surprisingly often. ## The four costs that go missing Model and infrastructure spend per decision, which at real volume is a monthly line rather than a rounding error. The time of whoever handles escalated exceptions. This is a designed cost, not a defect, and it needs an honest escalation-rate estimate rather than an assumption of nearly zero. Maintenance when the world moves. A supplier changes a form, a channel changes an interface, and somebody has to notice and fix it. The process owner's continuing time after handover, because an automation nobody owns degrades quietly while the dashboards keep looking fine. Ask for all four as one monthly figure, add twelve of them to the build price, and compare those totals. The [four questions that break most ROI numbers](/en/blog/roi-math-for-automation-projects) then have something to work on, because payback at three to six months cannot be checked at all without a monthly cost. ## Two structural things worth checking The payment schedule. If every milestone is a date rather than a working thing, you are funding elapsed time. At least one payment should be tied to something running in production that you can look at. The timeline against the scope. We put one to three workflows into production in two to four weeks, so a quote for a single workflow measured in months is describing different work — possibly a replacement project, possibly a discovery exercise with a build attached. Neither is wrong, but you should know which one you are buying. ## The best question to ask What am I not getting for this. A good vendor answers immediately and specifically, because they have already decided what is out of scope: the edge cases that stay manual, the system not integrated in this phase, the report not included. Ours should tell you which parts stay with a person by design, since that boundary is deliberate rather than a limitation. A vendor who cannot answer has either not thought about scope or is postponing the conversation until it becomes a change request. Both produce the same argument in month three. ## Before any of this arrives The quote is easier to read when you already know the answers. Count your own volume from your own systems, name the owner, and know which of your systems is authoritative for each disputed field. That is the same preparation described in [what makes a process map worth having](/en/blog/protocol-break-mapping-the-process), and it converts a quote review from an exercise in trust into an exercise in arithmetic. If the numbers in the quote disagree with yours, you have a specific conversation rather than a general unease. And if the readiness conditions in [when not to automate yet](/en/blog/what-size-company-should-not-automate-yet) are not met, the best-written quote in the world is still a quote for the wrong project. ## FAQ ### What should be itemised, at minimum? Six things, and any of them missing is a question rather than a deal-breaker. The measurement or discovery stage, priced separately, because a quote produced without one is a guess about a process nobody has counted. The build itself, tied to a named process rather than to a technology. Integration per system, listed one by one, since integrations are where estimates go wrong and a single combined figure hides which one is the risk. The monthly running cost, which is the line most often absent. Handover, including what documentation exists at the end and who receives it. And support terms after go-live, with a response time attached to them. A quote with all six can be compared to another quote with all six, which is the only way the price comparison means anything at all. ### What does a single-line total actually hide? Which assumptions can move, which is the thing you most need to see. Automation costs are dominated by volume and by integration surface, and both are estimates rather than facts at quoting time. When they are itemised you can ask what happens at half the assumed volume, or what the price becomes if the fourth system turns out to need a different approach, and the answers tell you where the risk sits. When they are collapsed into one number, every later change becomes a renegotiation from a position where you have no idea which part moved. The vendor is not necessarily hiding anything; single-line quotes are often just lazy. But the effect is identical, and the request to break it out costs nothing and is refused surprisingly often. ### Which running costs get left out most often? Four, in our experience, and together they are the difference between a project that pays back in four months and one that pays back in eleven. Model and infrastructure spend per decision, which at real volume is a monthly line rather than a rounding error. The time of whoever handles escalated exceptions, which is a designed cost rather than a defect, and which requires an honest estimate of the escalation rate instead of an assumption that it is near zero. Maintenance when the surrounding world moves, because a supplier changes a form and a channel changes an interface and somebody has to notice and fix it. And the process owner's continuing time after handover, since an automation nobody owns degrades quietly while the reporting keeps looking fine. Ask for these as one monthly figure and add twelve of them to the build price before comparing anything. ### What is the single best question to ask about a quote? What am I not getting for this. A good vendor answers immediately and specifically, because they have already decided what is out of scope and why: the edge cases that stay manual, the system that will not be integrated in this phase, the report that is not included. A vendor who cannot answer has either not thought about scope or is avoiding the conversation until it becomes a change request, and both produce the same argument in month three. The follow-up worth asking is what happens when the assumptions turn out to be wrong, which is a question about how change is priced rather than about whether change will happen. It always happens. Quotes differ in whether they admit it in advance. --- # AI Visibility for Operators, Measured on Our Own Site URL: https://inite.ai/en/blog/aeo-for-operators-what-it-buys-you Date: 2026-08-23 Author: Mikhail Savchenko Category: AEO Tags: AEO, AI Visibility, Operations, Strategy ## Direct Answer Being readable by AI engines and being recommended by them are different achievements, and only the first is a technical problem. Our own site scores 90 out of 100 on AI readiness and 25 out of 100 on AI visibility: the crawlers can parse everything and the engines still rarely mention us. Brand recall reaches 2 of 6 engines and category recommendations 0 of 6, which is the number that matters, because a buyer asking an assistant for a vendor is asking a category question. Closing that gap is earned through evidence and citations rather than through more schema markup. ## Key Facts - Our own site scores 90 out of 100 on AI readiness and 25 out of 100 on AI visibility. - Brand recall reaches 2 of 6 engines, and category recommendations reach 0 of 6. - Search Console recorded 30 clicks and 3,262 impressions across a month, with zero impressions on any commercial query. - llms.txt appeared on 10.13% of domains with no measurable citation lift in SE Ranking's November 2025 study. - Google's AI Mode runs at roughly 93% zero-click. ## Our own numbers, since they make the point better An AI visibility audit of inite.ai returns 90 out of 100 for readiness and 25 out of 100 for visibility. That means the crawlers can parse everything we publish and the engines still rarely mention us. Two of six engines recall the brand when handed the name. Zero of six recommend us in category, which is the question a buyer actually asks. Search Console tells the same story from the other side: 30 clicks and 3,262 impressions across a month, and zero impressions against any commercial query. Not low positions. Absence. We are publishing this because the alternative is writing about AI visibility from behind a number we have not earned, and because the gap itself is the useful thing to understand. ## Readiness is cheap. Visibility is not. | | AI readiness | AI visibility | | --- | --- | --- | | What it measures | Whether an engine can read you | Whether it chooses to mention you | | Who fixes it | A developer, in a fortnight | Evidence accumulated over months | | Improves | Immediately | Slowly, if at all | | Cost | Low and one-off | Ongoing | | What it is worth alone | Nothing | The whole thing | The two get priced as if they were the same product. A vendor who sells visibility and delivers readiness has delivered something real, much cheaper than what you thought you bought, and you will not notice for a quarter because the readiness score moves immediately. ## The number to watch Category recommendation, not brand recall. Brand recall asks whether an engine knows you exist once your name is supplied. It is easy to improve and worth very little, because someone who already knows your name has other ways to reach you. Category recommendation asks whether you appear when somebody describes a problem and asks who solves it. That is the question a buyer types. Our own split of 2 of 6 against 0 of 6 is the difference between being findable and being recommended, and only the second has revenue attached. Ask any vendor which of those two their headline number describes. The answer is informative. ## What actually seems to move it Nothing exotic, and nothing that can be finished in a fortnight. Answer a specific question completely, on a page that is about that question, in the form the question is asked. Mark it up so the answer is extractable rather than buried in a narrative. Then get corroborated somewhere that is not your own domain, because an engine weighing whether to recommend a vendor is looking for something that is not the vendor's own claim about itself. What does not appear to move it is the thing most commonly sold. llms.txt appeared on 10.13% of domains and SE Ranking's November 2025 study across 300,000 domains could not detect a citation lift attributable to it. We publish one anyway because it is inexpensive to maintain. That is a much weaker claim than the one usually made for it, and [the full argument is here](/en/blog/is-llms-txt-dead-2026). ## Why bother, when 93% of AI Mode is zero-click Because the traffic is not the asset. In a zero-click answer the assistant states a conclusion and names sources. Being one of those names puts you in the shortlist a buyer carries into their next conversation, including the one where they eventually type your name directly. Treating it as a traffic channel produces the wrong measurement and then the wrong conclusion, which is almost always that it did not work. Count whether you are named in answers to the questions your buyers ask, and watch branded search over the following months. If your reporting counts only sessions, a successful program and a failed one look identical. ## What an operator should do about it Three things, in order, and none of them is buying a tool. [Find out what an assistant currently says](/en/analyze) when asked who does what you do in your city or sector. That takes an afternoon and it is the only baseline that matters. Fix readiness once, because it is cheap and because being unreadable makes everything after it pointless. Whether you let the crawlers in at all is a business decision rather than a technical one, and [the allowlist post](/en/blog/ai-crawler-allowlist-2026) sets out the trade. Then spend the ongoing effort on evidence rather than markup: cases with numbers, answers to real questions, and corroboration on domains you do not own. That is slow, and it is the part that separates the two scores. The mechanics, with the schema that matters and the part that does not, are in [the AEO guide](/en/blog/aeo-complete-guide-2026). ## The honest summary We are good at the cheap half and bad at the expensive half, and we can prove both with numbers. Anyone selling you the cheap half at the price of the expensive one will show you a score that improves in week two. Ask what it measured. ## FAQ ### What is the difference between AI readiness and AI visibility? Readiness is whether an engine can read you. Visibility is whether it chooses to mention you. The first is a technical checklist that a competent developer can complete in a fortnight: clean markup, structured data, fast pages, a sensible robots policy, content that answers questions in the form questions are asked. The second is a reputation outcome that depends on whether anything outside your own domain corroborates what you say about yourself. This distinction matters commercially because the two are priced as if they were the same product. A vendor selling AI visibility and delivering readiness has delivered something real and much cheaper than what you thought you bought, and you will not discover the difference for a quarter, because a readiness score improves immediately and a visibility score does not. ### Which number should a small business actually watch? Category recommendations, not brand recall. Brand recall asks whether an engine knows you exist when your name is given to it, and it is easy to improve and worth very little, because a customer who already knows your name has other ways to reach you. Category recommendation asks whether you are among the answers when somebody describes their problem and asks who solves it, which is the question an actual buyer asks an assistant. On our own site the split is stark: 2 of 6 engines recall the brand, and 0 of 6 recommend it in category. That second zero is the commercially meaningful one, and it is the only one worth reporting to anyone who is paying. Ask any vendor which of the two their number describes. ### Does llms.txt help? There is no measured evidence that it does, and we publish one anyway for a narrower reason. SE Ranking's November 2025 study across 300,000 domains found it on 10.13% of domains and could not detect a citation lift attributable to it. That is not proof it is useless, but it does mean that anyone selling it as a visibility fix is selling ahead of the evidence. Content still has to answer a specific question completely and be corroborated somewhere that is not your own site. We keep llms.txt because it costs little to maintain, which is a much weaker claim than the one usually made for it. ### If zero-click is 93%, why bother appearing at all? Because the remaining traffic is not the point. In a zero-click answer the assistant states a conclusion and names its sources, and being one of those named sources is a different asset from a visit: it puts you in the shortlist a buyer takes into their next conversation, including the one where they eventually type your name directly. Treating this as a traffic channel produces the wrong measurements and then the wrong conclusion, which is usually that it did not work. Measure whether you are named in answers to the questions your buyers ask, and measure branded search volume over the following months, because that is where the effect surfaces. If your reporting only counts sessions, a successful AI visibility program is indistinguishable from a failed one. --- # Why Your Last Chatbot Failed, and It Was Not the Model URL: https://inite.ai/en/blog/why-your-last-chatbot-failed Date: 2026-08-22 Author: Olga Fedotova Category: Comparison Tags: Comparison, Operations, Customer Support, Automation ## Direct Answer Most failed chatbots were not failures of the model. They were scoped to answer everything instead of a defined set of things, given no access to the systems that hold the real answers, and measured on deflection rate, which pays a bot to avoid handing a customer to a person. That last one is the root cause of the experience your customers hated: the metric rewarded exactly the behavior that made it infuriating. A support automation worth deploying answers what it can verify, hands over the rest immediately with the full conversation attached, and is measured on resolution and on how fast the handover happens. ## Key Facts - Deflection rate sits on most chatbot dashboards, and it rewards exactly 1 behavior: not handing over. - In a brokerage deployment where the automation answered only what the listing could verify, response time fell from 6 hours to 8 minutes. - We put 1-3 workflows into production in 2-4 weeks, and support is often not the one we recommend first. - The entry point is a 15-minute diagnostic, free, before anyone talks price or scope. ## The metric caused the experience Most chatbot dashboards lead with deflection rate: the share of conversations that ended without a human. Read that definition again from the customer's side. Every handover counts as a failure. Every refusal to hand over counts as a win. A customer who gave up and closed the window scores identically to a customer who was helped. A system optimized against that number learns to keep people circling through restated questions and suggested articles. This is not a subtle misalignment. It is the precise mechanism behind the experience people are describing when they say they hate chatbots, and it was designed in on purpose, by whoever chose the metric. ## Three other things that were probably true | What went wrong | What it looked like to the customer | Fixable by | | --- | --- | --- | | Scoped to answer everything | Confident wrong answers on edge cases | Defining what it may answer | | No access to your systems | Paraphrasing the FAQ page | A read path into orders, stock, billing, calendar | | Nobody read the transcripts | The same failure every week for a year | One person, one hour, weekly | The second is the one that quietly determines everything. [A support automation connected to nothing](/en/automation/customer-support) can only restate published content, so it competes with your own search box and loses, because the customer read that page before opening the chat. The questions that generate contacts are specific and personal. Where is my order. Is this still available. Why was I charged this. Can I move my appointment. Answering any of them needs a read path into a real system, and building that path is most of the actual work. A proposal that skips it is quoting for a wrapper around a knowledge base. The demo will look excellent, because demos ask general questions. ## What a working version does differently It answers only what it can verify, and it says where the answer came from. In a brokerage deployment we ran, the automation answered what the listing itself could answer: floor, area, price, what is included, whether the property is still available. Everything else went to a named agent with the conversation attached. Response time went from 6 hours to 8 minutes, and the reason it worked is that the automation never guessed. Anything binding goes to a person by design. Pricing outside the published rate, terms, commitments. That boundary is the subject of [our rules for keeping a person in the loop](/en/blog/safe-ai-framework-human-in-loop), and it is not a limitation we apologize for: an automation that agrees something on your behalf at 2am is a liability rather than a feature. ## Handover is the whole product Three things go wrong here and all three are cheap to fix. The escape hatch is hidden, so the customer has to guess a magic phrase to reach a person. Say in the first message that a human is available. The handover arrives as a bare alert, so the agent opens by asking what the customer has already explained twice. Carry the transcript across, or the automation has cost time rather than saved it. The queue behind the handover is not staffed for what the bot escalates, so a fast refusal becomes a long silence. That is a capacity decision, and it has to be made before launch rather than discovered in week two. Route on the first sign of frustration, not the third. The cost of an unnecessary handover is a few minutes of an agent's time. The cost of a refused one is the customer. ## When we say do not build it More often in this category than in any other, and usually for one of three reasons. The contacts are mostly things a bot cannot verify. The volume is too low for anyone to maintain it. Or the real problem is that human response is slow, and support automation would be a decoration over that. The third case deserves naming because it is common. If inquiries wait four hours because nobody is there at seven in the evening, a bot that says something friendly and unhelpful at seven in the evening has not fixed the wait, it has automated it. The money is better spent on routing, on coverage, or on removing the reason people are contacting you at all. That is the same test as the [fifth readiness condition](/en/blog/what-size-company-should-not-automate-yet): if the bottleneck is not here, making this part faster changes nothing anyone can bank. ## What to ask the next vendor Which systems will it read from, and what will it do when that read fails. What is it measured on, and if the answer is deflection rate, what happens to the number when it hands over correctly. How does a customer reach a person, in how many messages, and who is waiting when they arrive. And ask to see a transcript from a real deployment on a bad day. The [order-processing split between rules and model](/en/blog/order-processing-equipment-rental) is what a defensible answer to the first question looks like: deterministic questions answered by rules, unstructured input read by a model, and the two never swapped. ## FAQ ### What is actually wrong with measuring deflection rate? It pays the system to do the thing customers hate. Deflection rate counts conversations that ended without a human, so every handover is scored as a failure and every refusal to hand over is scored as a win, regardless of whether the customer got what they came for. A bot optimized against that number learns to keep people in the loop of restated questions and suggested articles, because a customer who gives up and closes the window counts identically to a customer who was helped. That is not a subtle misalignment; it is the exact mechanism behind the experience most people are describing when they say they hate chatbots. Measure resolution instead, and measure time to handover, and the same technology produces an entirely different experience because the incentive now points the same way the customer does. ### Our bot could only repeat the FAQ page. Why? Because it had no access to the systems that hold the answers, which is a scoping and integration decision rather than a model limitation. A support automation connected to nothing can only paraphrase published content, so it is competing with your own search box and losing, because the customer has usually already read that page before opening a chat. The questions that actually generate contacts are specific and personal: where is my order, is this still available, why was I charged this, can I move my appointment. Answering those requires a read path into the order, inventory, billing or calendar system, and building that path is most of the real work. Any proposal that skips it is quoting for a wrapper around a knowledge base, and the demo will look excellent because demos ask general questions. ### How should the handover actually work? Immediately, visibly, and with the whole conversation attached rather than as a fresh ticket. Three things go wrong in practice and all three are fixable. The escape hatch is hidden, so the customer has to guess a magic phrase to reach a person, which converts mild irritation into anger. The handover arrives as a bare alert, so the agent opens with a question the customer has already answered twice and the automation has just cost time rather than saved it. And the queue behind the handover is not staffed for the volume the bot escalates, which turns a fast refusal into a long silence. Say plainly in the first message that a person is available, route on the first sign of frustration rather than the third, and carry the transcript across. ### When is the honest answer that you should not deploy one at all? When the contacts you receive are mostly things a bot cannot verify, when the volume is too low for anyone to maintain the thing, or when the real problem is that your human response is slow and support automation would be a decoration over that. The last case is common and worth naming: if inquiries wait four hours because there is nobody to answer them at seven in the evening, a bot that says something friendly and unhelpful at seven in the evening has not fixed anything, it has just made the wait feel automated. In that situation the money is better spent on routing, on coverage, or on removing the reason people are contacting you at all. We say no to support automation on these grounds more often than to any other category of project. --- # Your Developer Can Build It. That Is Not the Question URL: https://inite.ai/en/blog/in-house-developer-vs-agency-for-automation Date: 2026-08-21 Author: Anton Fenix Category: Comparison Tags: Comparison, Operations, Procurement, Automation ## Direct Answer Your developer can almost certainly build the workflow. The question is what stops being built instead, and whether a project with no deadline ever finishes. An internal build competes with the product roadmap rather than with a delivery date, which is why it routinely stretches while an external one lands in weeks. In-house genuinely wins on knowledge of your own strange systems and on long-term ownership; an agency wins on having watched this class of work fail before. The arrangement that usually beats both is an external build with internal ownership from day one. ## Key Facts - We put 1-3 workflows into production in 2-4 weeks against a delivery date, which an internal project rarely has. - Payback lands at 3-6 months, and that clock only starts when the build actually finishes. - A single-developer internal build is 1 person deep, and the exception queue outlives most tenures. - A first delivery is 1-3 workflows in production in 2-4 weeks, handed to the client's own team with documentation and monitoring. ## The capability question is the easy one Your developer can build it. In most cases that is simply true, and any comparison that opens by casting doubt on it should be treated with suspicion. The interesting questions are different. What stops being built instead, and does a project with no deadline ever get finished. ## The invisible price An internal build costs whatever was next on the roadmap. That price never appears as a line item, which is why it rarely appears in the decision either. If the thing your developer would otherwise ship is the product your customers pay for, the automation is expensive in a way the invoice will never show. If your team is genuinely under-loaded, the arithmetic flips and building internally is plainly right. The useful move is to make it explicit. Name the feature or fix that will slip. Put a date on it. Show that date to whoever owns the roadmap and see whether the trade still looks obvious. ## Why internal projects stretch They compete with a roadmap rather than with a deadline, and a project with no delivery date has no mechanism for being finished. The pattern is consistent enough to plan around. The build starts quickly and well. Then an urgent customer issue arrives, then a release, then somebody leaves, and the automation becomes the thing picked up between other things. Six weeks of work spread across eight months is not the same as six weeks of work. The operational problem stays unsolved for those eight months, and the requirements drift underneath the half-built system. External delivery is not faster because the people are better. It is faster because the work has a date, a fixed scope, and nothing else competing for the same hours. If you build internally, the fix is to give the project those three things rather than to hope. ## What each side actually wins | | In-house | Agency | | --- | --- | --- | | Knows your undocumented systems | Yes | Learns them, at a cost of days | | Has a delivery date | Rarely | By contract | | Has seen this fail before | Sometimes | This is what you are buying | | Still there in month four | Yes | Only if contracted | | Depends on one person staying | Usually | No | | Cost on the invoice | None | Real | | Cost to the roadmap | Real | None | The last two rows are the whole comparison, and they point in opposite directions. Everything else is a detail. ## The arrangement that usually beats both Not a choice. A division. External delivery against a date, internal ownership from the first week. The person who will own the workflow sits in the build rather than receiving a handover document at the end, so the knowledge transfers by participation instead of by paperwork. That is the model we work to, and it is why [the named-owner condition](/en/blog/what-size-company-should-not-automate-yet) belongs in the proposal rather than in the final week. It also produces the thing an internal build produces naturally and an external one often does not: somebody who genuinely understands why the system does what it does. ## On hiring for it Only if you have enough automation work to keep the person interested, and that bar is higher than it looks. One automation engineer is a single point of failure [in a way an agency is not](/en/compare). The work also has a retention problem nobody mentions: building the first three workflows is interesting, and maintaining them while the surrounding systems shift is not. Companies that hire for this and then run out of new things to build tend to lose the person inside a year and inherit a system only they understood. If the pipeline is real, a workflow a quarter indefinitely, hiring wins on economics by a wide margin. If it is two projects and then maintenance, buy the builds and keep the ownership. ## Before either Neither route helps if the process is not ready, and neither vendor nor employee should be quoting before somebody has counted. The [measurement week](/en/blog/rental-case-the-week-before) applies identically to an internal build, and an internal team is if anything more likely to skip the measurement because they already believe they know the process. They usually know their part of it. That is a different thing, and [why most process maps are useless](/en/blog/protocol-break-mapping-the-process) is about exactly that gap. ## FAQ ### Is it cheaper to build automation in-house? On the invoice, almost always. In total cost, it depends on something most comparisons leave out entirely, which is what your developer stops doing. An internal build has a real price equal to whatever was next on the roadmap, and that price is invisible because it never appears as a line item. If the thing they would otherwise be building is the product your customers pay for, the automation is far more expensive than it looks, even though no money changes hands. If your developers are genuinely under-loaded, the calculation flips and building internally is straightforwardly the right answer. The useful move is to make the opportunity cost explicit before deciding: name the feature or fix that will slip, put a date on it, and see whether the trade still looks obvious to whoever owns that roadmap. ### Why do internal automation projects take so much longer? Because they compete with a roadmap rather than with a deadline, and a project without a delivery date has no mechanism for being finished. The pattern is consistent enough to plan around: the build starts quickly and well, then an urgent customer issue arrives, then a release, then someone leaves, and the automation becomes the thing that gets picked up between other things. Six weeks of work spread across eight months is not the same as six weeks of work, because the operational problem it was meant to solve stays unsolved for those eight months and the requirements drift underneath it. External delivery is not faster because the people are better; it is faster because the work has a date, a defined scope and nothing else competing for the same hours. If you build internally, the fix is to give the project the same three things rather than to hope. ### What does an in-house developer genuinely do better? Two things, and both are worth real money. They know your systems, including the undocumented ones, the field that means something different from what its name suggests, and the integration that somebody wrote four years ago and nobody has touched since. An external team spends its first days discovering exactly that, and in an unusual estate those days can be a meaningful share of the project. The second advantage is permanence: they are still there in month four when a supplier changes a form, and their knowledge of the workflow compounds instead of leaving with a contract. This is why the strongest arrangement is usually not a choice between the two but a division: external delivery against a date, internal ownership from the first week, so the person who inherits it has been in the room the whole time rather than receiving a handover document. ### Should we hire someone specifically for automation work? Only if you have enough of it to keep them interested, and that bar is higher than it first appears. One automation engineer is a single point of failure in a way an agency is not, and the work has a retention problem that nobody warns you about: building the first three workflows is genuinely interesting, and maintaining them while the surrounding systems shift is not. Companies that hire for this and then have nothing new to build tend to lose the person within a year and inherit a system only they understood. If the pipeline is real — a workflow a quarter, indefinitely — hiring is the better economics by a wide margin. If it is one or two projects and then maintenance, buy the builds and keep the ownership, which costs a fraction of a salary and does not depend on one person staying. --- # Why Most Process Maps Are Useless, and What Fixes Them URL: https://inite.ai/en/blog/protocol-break-mapping-the-process Date: 2026-08-20 Author: Mikhail Savchenko Category: Methodology Tags: Methodology, Operations, Process Audit, INITE Protocol ## Direct Answer Most process maps are drawn from interviews, which means they describe the process as designed rather than as run, and the gap between those two is where the losses live. A map worth having carries a number at every handoff, includes the workarounds people actually use, and is built partly from system data rather than entirely from what anyone says. In one Break stage we mapped 3 candidate workflows this way and priced the hours they lost at $3,140 a week, which is the figure every later ROI claim gets checked against. ## Key Facts - In one Break stage we mapped 3 candidate workflows and priced the hours lost across them at $3,140 a week, or $151K a year over 48 working weeks. - One of those workflows took a median 38 hours to first response, and 24% of its leads got no reply inside five working days. - Invoice and time-entry reconciliation on the same engagement ran 8 hours a week at a 22% error rate. - We shadow the real work for 4-8 hours per target workflow and ingest 90 days of operational data. - The Break stage runs in weeks 1-2 of the engagement and can end it, with the diagnostic deposit refunded. ## The map most companies already have Almost every operation has a process diagram somewhere. Boxes, arrows, a swimlane per department, drawn at some point by someone who interviewed everybody. It is usually accurate and almost always useless, for one reason: it describes the process as designed. The work is done differently, and the difference is the entire subject. ## Two things make a map worth drawing The first is a number at every handoff. Throughput, error rate, cycle time. Without them a diagram tells you the steps exist and nothing about which one costs anything, so every decision about what to fix gets made on impression. With them, the steps sort themselves into the annoying and the expensive, and those two groups overlap much less than anyone expects. The second is the real route, including the workarounds. If three people rely on a shared spreadsheet that appears nowhere in the official process, that spreadsheet belongs on the map. [It is load-bearing, and it is usually where the contradictions that break automation](/en/protocol) are hiding. ## Where the numbers come from Interviews plus system data, and the two answer different questions. People are accurate about their own steps and unreliable about waiting, frequency and exceptions. Somebody describing their part of a process gives a good account of what they do and a poor one of how long the work sits before it reaches them, because nobody experiences a queue they are not standing in. So we shadow the real work for four to eight hours per target workflow and ingest ninety days of operational data from whatever systems already hold it. The data corrects the distortions: it shows the distribution rather than the impression, covers the nights and weekends nobody recalls, and includes the cases that were abandoned, which by definition nobody remembers. The interviews then become useful for a different job, which is explaining why the data looks the way it does. ## What comes out | Artifact | What it contains | What it is for | | --- | --- | --- | | Process map | Every step, with throughput, error rate, cycle time at each handoff | Locating the expensive steps rather than the irritating ones | | Cost-of-chaos report | Money lost per week to rework, missed handoffs and waiting | The baseline every later ROI claim is checked against | | Priority matrix | Each candidate workflow scored on feasibility and return | Deciding what gets built and what is explicitly deferred | On one engagement, mapping three candidate workflows priced the hours they lost at $3,140 a week. The largest single line was project status updates: 26 active engagements, one note a week each, 40 minutes of consultant time per note, at $95 an hour loaded. Another, invoice and time-entry reconciliation, ran 8 hours a week at a 22% error rate. Those numbers are the point. Not because they are large - they are not, and a firm of sixty people can carry a loss that size for years without anyone flinching - but because the later claim about improvement has something specific to be measured against, and nobody can quietly compare an after number to an imagined before. What the total deliberately left out matters as much. The same engagement had a median first response of 38 hours on inbound leads, and 24% of those leads got no reply inside five working days. Both are almost certainly worth more than every hour in the total. Neither is in it, because pricing them means assuming a conversion rate and a deal value, and an assumption buried inside a baseline makes every comparison built on top of it unfalsifiable. ## The disagreement is the finding The most useful thing a map produces is rarely on the map. Three people describe the same process and the descriptions differ. That is not sloppiness; it is information. The places where accounts diverge are almost always where the exceptions live, where a workaround has replaced the official route, or where two systems disagree about a fact and different people have picked different winners. That last one decides the size of the eventual project more than any technology choice, which is the argument made at length in [the decision that shaped a three-week build](/en/blog/rental-case-the-decision-that-shaped-it). ## What happens if the numbers say no The stage ends the engagement and the diagnostic deposit is refunded. That has to be a real outcome or none of the measurement means anything, and it is the same reason the [readiness conditions](/en/blog/what-size-company-should-not-automate-yet) are worth applying before a proposal rather than after one. The client still leaves with the map, the baseline and the matrix. All three are useful whether or not anything gets built, and they make any future project cheaper, because the expensive part of automation is finding out how the work actually moves rather than writing the code that moves it. A diagnostic that produces nothing reusable has produced a sales document rather than an audit. That distinction is worth applying to us as much as to anyone, and [the measurement week at a rental firm](/en/blog/rental-case-the-week-before) is what it looks like when it is done properly on a real operation. ## FAQ ### What makes a process map worth the time it takes? Numbers at the handoffs, and honesty about the paths people actually take. A diagram of boxes and arrows with no quantities is a picture of an org chart pretending to be an analysis: it tells you the steps exist without telling you which one is costing anything, so every subsequent decision about what to fix is made on impression. A map worth having marks throughput, error rate and cycle time at each handoff, which immediately sorts the steps into the ones that are annoying and the ones that are expensive, and those two groups overlap far less than people expect. The second requirement is that it shows the real route including the workarounds. If three people use a shared spreadsheet that appears nowhere in the official process, the spreadsheet belongs on the map, because it is load-bearing and because it is usually where the contradictions that break automation are hiding. ### Why not just interview the people who do the work? Interview them, but do not stop there, because people are accurate about their own steps and unreliable about waiting, frequency and exceptions. Someone describing their part of a process gives you a good account of what they do and a poor account of how long the work sits between their part and the next one, because nobody experiences the queue they are not in. They also describe the process as it is supposed to run, which is not dishonesty but the natural way anyone answers a question about their job. Ninety days of system data corrects both distortions: it shows the distribution rather than the impression, covers the nights and weekends nobody recalls, and includes the cases that were abandoned, which by definition nobody can remember. The interviews then become useful for a different purpose, which is finding out why the data looks the way it does. ### What is the cost of chaos number, and how is it calculated? It is the money lost per week to manual rework, missed handoffs and waiting, calculated across the candidate workflows rather than across the company. On one engagement that figure was $3,140 a week over three workflows, and its purpose is narrow but important: it is the baseline that every later return claim gets measured against, so that nobody can quietly compare an after number to an imagined before. It is deliberately built from loaded hourly costs rather than headline salaries, from the low month rather than the good one, and from volumes taken out of the systems rather than out of a conversation. It is also deliberately incomplete: a loss that cannot be priced without assuming a conversion rate stays on the page as a finding and out of the total, because an assumption inside a baseline quietly makes every later comparison unfalsifiable. A cost-of-chaos number produced any other way flatters the project, and a flattered project fails the same arithmetic later, only after somebody has been paid. ### What does the client keep when the engagement ends? The map, the baseline and the priority matrix, and all three are useful whether or not anything gets built. This matters more than it sounds, because it is what makes it possible for the Break stage to conclude that there is no project worth doing. If the arithmetic does not survive, the engagement ends there and the diagnostic deposit is refunded, and the client still leaves holding a quantified picture of their own operation that they did not have before. It also makes the second project cheaper, since the expensive part of any automation is finding out how the work actually moves rather than writing the code that moves it. A vendor whose diagnostic produces nothing reusable has produced a sales document, not an audit. --- # The Decision That Shaped a Three-Week Rental Build URL: https://inite.ai/en/blog/rental-case-the-decision-that-shaped-it Date: 2026-08-19 Author: Mikhail Savchenko Category: Case Study Tags: Case Study, Operations, Equipment Rental, Methodology ## Direct Answer The measurement had shown that machines physically standing in the yard were being reported as unavailable, which pointed at the data rather than at the staff. That left three options: replace the rental system, put a language model on top of the existing spreadsheets, or make one record authoritative and give an asset more than one possible state. The third was chosen, and it is why the build took 3 weeks rather than months. It also had a cost nobody puts in a proposal: somebody had to give up their own spreadsheet. ## Key Facts - 3 options were on the table, and 2 of them were measured in months rather than weeks. - The chosen design replaced 1 boolean availability field with 7 asset states. - It left the language model exactly 1 job, reading free-text inquiries, instead of the availability question. - The build took 3 weeks and moved booking-to-dispatch from 4 hours to 3 minutes. - Peak-season capacity rose 2.5x afterwards with no change to headcount. ## What the measurement left us with The week of counting had produced one finding that changed the project: machines physically standing in the yard were being reported as unavailable, often enough to explain both the double-bookings and a share of the refusals. That is a sentence about data rather than about people. Had the count come out near zero, the honest conclusion would have been that staff were overloaded and the answer was capacity or routing. It did not, so the answer was somewhere in how availability was recorded. Three options followed. Two of them were months. ## The three options | Option | Time | Why it was rejected or chosen | | --- | --- | --- | | Replace the rental system | Months | Fixes one field by migrating everything, during the season | | Put a model over the spreadsheets | Weeks | Answers the wrong data faster and less traceably | | Make one record authoritative, give assets real states | 3 weeks | Fixes the specific thing that was wrong | The first is the option most often proposed when a data model is at fault, and it is usually the wrong scale of response. A replacement migrates history, retrains everyone, and runs two half-trusted systems in parallel, all to correct one field. In a seasonal business, doing that during the months when the operation is already straining is not a detail. The second is the option that would have been easiest to sell. A model reading the existing sheets and answering availability questions demos well and fails in the exact way the measurement had already predicted: the data says a machine is unavailable while it stands in the yard, and the model repeats that with more confidence and less traceability than the spreadsheet had. ## Why the model did not get that job There is a general rule underneath this particular choice, and it is worth separating from the specifics. An availability check has to return the same answer to the same question every time, and it has to be explainable afterwards when a customer asks why they were refused. A probabilistic answer to a deterministic question is a defect, and no improvement in the model changes that. So the model got exactly one job in the finished system: reading inquiries that arrive as free text at eleven at night, naming a machine in the customer's own words with dates in a format no field expects. That is genuinely hard for a rule and genuinely easy for a model. The split is set out in full in [where the four hours go](/en/blog/order-processing-equipment-rental). ## What the chosen design changed Two things, and the second mattered more. Availability stopped being one true-or-false field and became seven explicit states, so a machine returned but awaiting inspection, a machine in transit, and a machine reserved but unconfirmed all stopped reporting as free. Then one record became authoritative about what is reserved, and the check moved to the moment of commitment rather than the moment of inquiry. That second change is what ended the double-bookings, because the conflicts had been living in the window between somebody looking at a shared sheet and somebody promising a machine. Neither change is exotic. Neither needed new technology. Both needed a decision that somebody had to make and nobody had made. ## The cost that appears in no proposal Making one record authoritative means somebody has to stop keeping their own copy. Every operation of this kind has one or two people running a private spreadsheet, and they are not being obstructive. They started keeping it because at some point the official system was wrong and their sheet was right, and it has been quietly holding the business together ever since. Asking them to trust a system that previously let them down is the real work, and it lands very differently depending on whether their objection was listened to first. This is the part that most nearly derailed the project. It appears in no proposal we have ever seen, including our own earlier ones, and it is why the source-of-truth decision is now named explicitly at the start rather than discovered in week two. ## What it produced Booking-to-dispatch went from 4 hours to 3 minutes. Double-bookings were eliminated. Peak-season capacity rose 2.5x with the team that was already there, and the build fit in 3 weeks. The full set is on the [case page](/en/cases/equipment-rental-automation). The three weeks are a consequence of the decision rather than of speed. The two rejected options were not slower versions of the same project; they were different projects, and one of them would have been finished around the time the season ended. ## The transferable part Before agreeing to anything, find the fact that two of your systems disagree about, and decide which one wins. That decision costs nothing, takes an afternoon, and determines the size of every project that follows it. The measurement that surfaces it is described in [the week before](/en/blog/rental-case-the-week-before), and it is the same reason [contradictory data is the one kind you should not automate around](/en/blog/what-size-company-should-not-automate-yet) until the question has an answer. ## FAQ ### Why not replace the rental software, since the data model was the problem? Because the data model was wrong in one specific way rather than wrong throughout, and replacing a system to fix one field is a months-long project carrying risks that have nothing to do with the original problem. A replacement means migrating history, retraining everybody, and running two systems in parallel for a period during which both are half-trusted. It also puts the whole operation on a new tool during the exact months when it is already struggling, and in a seasonal business the timing of that is not a detail. The narrower fix was to leave the existing system holding what it held correctly and to make one record authoritative about the single thing it was getting wrong, which is reservations against assets. That is weeks of work instead of months, and it leaves the option of replacing the system later open rather than spending it now. ### What was wrong with putting a language model over the existing spreadsheets? It would have answered the availability question faster and just as wrongly, which is a worse outcome than answering it slowly. The measurement had already established that the underlying data reported machines as unavailable while they stood in the yard, and a model reading that data reproduces the error with more confidence and less traceability. There is also a subtler problem worth naming, because it recurs in most projects where a model is proposed as a shortcut: an availability check must return the same answer to the same question every time, and must be explainable afterwards when a customer asks why they were refused. A probabilistic answer to a deterministic question is a defect regardless of how good the model is. The model earned a place in the finished system, but at the front door on unstructured inquiries, which is work a rule genuinely cannot do. ### What did the chosen design actually change? Two things, and the second is the one that mattered. First, availability stopped being a single true-or-false field and became an explicit set of states, so that a machine returned but awaiting inspection, a machine in transit, and a machine reserved but unconfirmed all stopped reporting as available. Second, one record became authoritative about what is reserved, and the availability check moved to the moment of commitment rather than the moment of inquiry. The second change is what eliminated double-bookings, because the conflicts were living in the window between someone looking at a shared sheet and someone promising a machine. Neither change is exotic and neither required new technology. They required a decision that somebody had to make and that nobody had made, which is the usual shape of these projects. ### What did that decision cost? Somebody had to stop keeping their own copy, and that is a political cost rather than a technical one. In practice every operation of this kind has one or two people running a private spreadsheet, usually because at some point the official system was wrong and their sheet was right. Those people are not being obstructive; they are holding the workaround that has been keeping the business functional. Making one record authoritative means asking them to trust a system that has previously let them down, and that request lands very differently depending on whether their objection has been listened to first. This is the part of the project that most nearly derailed it, it appears in no proposal we have ever seen including our own early ones, and it is why we now name the source-of-truth decision explicitly at the start rather than discovering it in week two. --- # Six Hours to Eight Minutes: Inquiries at a Brokerage URL: https://inite.ai/en/blog/lead-response-real-estate-agency Date: 2026-08-18 Author: Mikhail Savchenko Category: Automation Tags: Automation, Operations, Real Estate, Lead Response ## Direct Answer At a brokerage, response time is set by routing rather than by typing. An inquiry arrives naming a specific property, and someone has to decide which agent owns it, whether this person has already inquired through another portal, and whether the listing is even still available. Those three decisions are where the hours go, and all three are rules rather than judgment. Automating them took one brokerage from a 6 hour response time to 8 minutes and shortened the deal cycle from 14 days to 5, in a build that took 4 weeks. ## Key Facts - At one brokerage, lead response fell from 6 hours to 8 minutes and the deal cycle from 14 days to 5. - Document preparation at the same agency went from 2 days to 20 minutes, and agent productivity rose 60%. - That build took 4 weeks, inside our usual 2-4 week window for 1-3 workflows. - A buyer inquiring through 3 portals about the same flat is 1 lead, and treating it as 3 is how agencies contact people twice. - A first delivery is 1-3 workflows in production in 2-4 weeks, handed to the client's own team with documentation and monitoring. ## An inquiry is not one thing A buyer messages about a two-bedroom flat. That single message is three separate questions, and a brokerage answers them in sequence, usually with a human waiting between each. Who is this person, and have we already spoken to them? They may exist in the pipeline under a different number from a different portal, because a serious buyer inquires in several places on the same evening. Which listing is this, and is it still available? Under offer since Friday is a different conversation, and an agent who does not know that is about to waste an afternoon. Which agent owns it? Area, specialization, current load, and who is actually working this weekend. None of those three is judgment. All three are rules, and the hours between an inquiry arriving and a reply going out are almost entirely the wait for a person to apply them. ## Where the six hours went At the brokerage we worked with, leads lived in agents' personal notebooks and listing updates went out by manual email. Nobody had a pipeline view, so a stalled deal was invisible until someone remembered to ask. | Step | Who did it | What it cost | | --- | --- | --- | | Notice the inquiry | Whoever checked that inbox | Minutes to hours, depending on the hour | | Check the buyer is new | The agent, from memory | Duplicate contact when memory failed | | Check availability | A call or a message to a colleague | The wait for the colleague | | Decide the owner | Whoever was in the office | Uneven load, best agent buried | | Reply | The agent | Minutes | The reply itself was always minutes. Everything above it was the six hours. ## Deduplication is the unglamorous half A buyer who inquires through three portals about the same flat is one lead. Treating that as three is how agencies phone the same person twice in an evening and look disorganised at exactly the moment they are being compared to two competitors. Matching on phone number alone fails, because portals mask numbers. Matching on name alone fails, because names repeat. What works is the combination of contact detail, listing and time window, which is a rule rather than a model, and which no salesperson can apply reliably at nine in the evening. This is the least interesting part of the project to describe and one of the most valuable in practice. ## The routing rule is a business decision Three rules are common and they optimize for different things. | Rule | Optimizes for | Quietly costs | | --- | --- | --- | | Round-robin | Fairness between agents | Conversion, when agents differ | | Specialization | Quality of the conversation | Uneven load, single points of failure | | Load-based | Everyone busy | Punishes the fastest agents with more work | Most agencies want a blend. The useful part of the project is often the conversation that forces the blend to be written down, because a rule nobody has stated is not being followed consistently by the humans either. We implement the rule the agency chooses. Choosing it is not our decision to make, and a vendor who arrives with an opinion about which of your agents deserves more leads has misunderstood the engagement. ## What answers at two in the morning Enough to hold the conversation, and nothing binding. It confirms the property is still available. It answers what the listing can answer: floor, area, price, what is included. It offers viewing slots from the assigned agent's actual calendar. It captures what the buyer is looking for, in their words. It does not negotiate, quote outside the published price, or commit to terms. Those reach the agent with the whole conversation attached, which is the difference between a handover and an alert. That boundary is the same one described in [our rules for keeping a person in the loop](/en/blog/safe-ai-framework-human-in-loop), and it is not a limitation we are apologizing for: an automation that agrees a price at 2am is a liability, not a feature. Most out-of-hours inquiries are lost because nobody confirmed the flat still existed, not because nobody negotiated overnight. By morning the buyer has two viewings booked elsewhere. ## Where the deal cycle actually came from Response time went from 6 hours to 8 minutes. The deal cycle went from 14 days to 5. These get quoted together, and the second is not caused by the first. What shortened the cycle was the rest of the same work: viewings booked without three phone calls, document packages prepared in 20 minutes instead of 2 days, and a pipeline everyone could see, so a stalled deal was visible while it was still recoverable. Agent productivity rose 60% on the same basis. A project that fixes only first response produces a beautiful first metric and a deal cycle that has barely moved. That distinction is worth insisting on when someone quotes you the dramatic number, and it is the same discipline as [the four questions that break most ROI claims](/en/blog/roi-math-for-automation-projects). ## Before committing to this Count two things from your own systems: how many inquiries arrive outside working hours, and how many are answered more than an hour after arrival. Then count how many buyers appear twice. Those three numbers size the whole project, and the method is the one in [the measurement week](/en/blog/rental-case-the-week-before). The full case is on the [brokerage case page](/en/cases/real-estate-deal-cycle), and what we deploy in this industry is set out at [AI automation for real estate](/en/industries/real-estate) and [lead response](/en/automation/lead-response). ## FAQ ### Why is a property inquiry harder to route than a normal sales lead? Because it is not one thing, it is three, and a generic lead router only handles the first. There is a person, who may already exist in your pipeline under a different phone number from a different portal. There is a specific listing, which has an owner, a status and possibly an exclusivity arrangement that decides who is allowed to handle it. And there is a routing rule, which in most agencies is a mixture of area, specialization, current load and who is actually working this weekend. Getting the person wrong means contacting someone twice and looking disorganised. Getting the listing wrong means an agent showing a property that went under offer on Friday. Getting the rule wrong means the agency's best performer is buried while a colleague sits idle. Ordinary CRM lead capture solves none of the three, which is why brokerages that have a CRM still answer in hours. ### What decides which agent gets the lead? That is a business decision and the honest answer is that we do not make it, we implement whichever one the agency has already made and often has never written down. There are three common rules and they optimize for different things. Round-robin is fair and simple and ignores that some agents close far better than others. Specialization by area or property type produces better conversations and concentrates load unevenly. Load-based routing keeps everyone busy and quietly punishes the agents who work fastest by giving them more. Most agencies want a blend, and the useful part of the project is usually the conversation that forces the blend to be stated explicitly, because an unstated rule cannot be automated and is also, in practice, not being followed consistently by humans either. ### What does the automation actually reply with at two in the morning? Enough to hold the conversation, and never anything binding. It confirms the property is still available, answers the questions that have factual answers from the listing itself such as floor, area, price and what is included, offers viewing slots from the assigned agent's real calendar, and captures what the buyer is actually looking for. What it does not do is negotiate, quote anything outside the published price, or commit to terms. Those go to the agent with the full conversation attached, which is the difference between a useful handover and an alert. The commercial point is narrow: most inquiries outside working hours are lost not because nobody negotiated overnight but because nobody confirmed the flat still existed, and by morning the buyer had booked two viewings with somebody else. ### How does faster response turn into a shorter deal cycle? Through fewer gaps, and it is worth being precise because the two numbers get quoted together as if one obviously causes the other. Response time falling from 6 hours to 8 minutes does not itself shorten a deal by nine days. What shortens it is that the same routing and pipeline work removes the other waits: the viewing booked without three phone calls, the document package prepared in 20 minutes instead of 2 days, and a pipeline everyone can see so that a stalled deal is visible while it is still recoverable. The response number is the one that gets attention because it is dramatic; the document and pipeline numbers are what actually move the cycle. A project that only fixes first response will produce a lovely first metric and a deal cycle that has barely moved. --- # What to Automate First, and Why It Is Not the Worst Job URL: https://inite.ai/en/blog/what-to-automate-first-in-a-small-company Date: 2026-08-17 Author: Mikhail Savchenko Category: Operations Tags: Operations, Procurement, Automation, Methodology ## Direct Answer Pick the first process by blast radius rather than by how much it is disliked. The most hated job is usually the one that touches contracts or money, and those are exactly the places where an early mistake costs a customer rather than a minute. A good first automation runs often enough to produce a signal within weeks, fails in a way somebody can undo, and has an owner who will notice. In practice that is most often inbound inquiry routing, which is why lead response tends to be first of the four processes we deploy most, and it is the one that teaches you the most about whether the next one is worth doing. ## Key Facts - We put 1-3 workflows into production in 2-4 weeks, so the first one is a choice among 3, not a question of whether. - Of the 4 processes we deploy most often, lead response is usually the one that goes first. - In a brokerage deployment, lead response fell from 6 hours to 8 minutes and the deal cycle from 14 days to 5. - Payback typically lands at 3-6 months, and a first project produces a usable signal in far less than that if it runs often enough. ## The instinct is wrong in a predictable way Ask a team [which process to automate first](/en/answers/is-my-process-worth-automating) and they will name the one they hate. That job is usually fiddly, high-stakes and infrequent. Contract preparation. The month-end reconciliation. The thing that has to be right and takes an afternoon and only happens twice a month. It is close to the worst possible first candidate, and for reasons that have nothing to do with whether it deserves to be automated eventually. ## Blast radius is the criterion The question that orders the list is what happens when the system gets it wrong, because on a first project it will. | Process | If it fails | Blast radius | | --- | --- | --- | | Inquiry routing | One reply goes to the wrong person | 1 conversation, recoverable in minutes | | Availability check | A booking is refused that could have been taken | 1 booking, recoverable same day | | Document assembly | A contract goes out with wrong terms | A customer, and possibly a legal problem | | Pricing or commitment | The company is bound to something | A customer, and the commitment stands | The first two are places to learn. The last two are places to be careful, and they are almost always where the complaints come from. ## Frequency is what turns a build into evidence The second criterion is how often it runs, and it decides how long you wait to know anything. A process that happens forty times a week gives you a usable answer in a fortnight. The same build on a process that happens twice a month tells you nothing for a quarter, and by then the team has stopped paying attention and the vendor has moved on. This is also the practical reason a first project should fit into two to four weeks. A long first project is a bet placed before any of the information has arrived, committing you to a supplier and a design at the moment you know the least. ## Which one it usually is Of the four processes we deploy most often, inbound inquiry handling is usually first. It runs constantly. A mistake is one misrouted reply. It touches few systems, so the integration work does not eat the schedule. And it produces a number quickly: in one brokerage deployment, response time went from 6 hours to 8 minutes, and the deal cycle from 14 days to 5. That second number is the one that funds the next project, and it is worth noticing where it came from. Faster replies did not save anyone's afternoon in a measurable way. They shortened the cycle, and shorter cycles convert better because fewer buyers cool off in the gap. ## When the obvious candidate touches money Split it rather than skipping it, because the part that touches money is rarely the part where the time goes. In an order flow, the availability check, the document assembly and the dispatch scheduling can all be automated while the commitment, the price and the terms stay with a person. You get the frequency and the saving without putting a three-week-old system in a position to bind the company. That is not a compromise for beginners. It is the design the finished system has anyway, for the reasons set out in [our rules for keeping a person in the loop](/en/blog/safe-ai-framework-human-in-loop), and [the rental order flow](/en/blog/order-processing-equipment-rental) is a worked example of exactly this split. ## What the first project is really for It is the cheapest chance you will get to learn three things no proposal can tell you. How your team reacts to a system making decisions, which is rarely how anybody predicts. How many exceptions the process actually produces once somebody is counting, which is almost always more than the estimate. And whether the vendor tells you about problems before you find them, which is the one that should decide whether there is a second project. Those answers change the shape of what comes next. Choosing a first project that cannot deliver them within a month is the expensive part of getting this wrong, and it is why the order matters more than the list. ## Before any of this None of the above helps if the process fails the readiness test in the first place, and the five conditions in [when not to automate yet](/en/blog/what-size-company-should-not-automate-yet) are worth running before choosing an order at all. Pick the volume from your own systems. Name the owner. Then take the frequent, recoverable one first, even though it is not the one anybody complains about. ## FAQ ### Why not start with the process the team complains about most? Because complaint tracks unpleasantness, and unpleasantness does not track either value or safety. The job everyone hates is usually the one that is fiddly, high-stakes and infrequent, which is close to the worst possible combination for a first automation. Infrequent means you wait months for enough runs to know whether it works. High-stakes means the first mistake is visible to a customer rather than to a colleague. Fiddly means it is full of exceptions, which is precisely the material that makes a build long and a result disappointing. There is a place for that process, and the place is second or third, after the team has learned how the system behaves and after somebody has watched an exception queue for a few weeks. Starting there is the most common way a first project sours an organization on the whole idea. ### What makes a good first candidate, concretely? Four properties, and the first two matter far more than the others. It has to run often, because frequency is what turns a build into evidence: a process that happens forty times a week tells you whether it works within a fortnight, while one that happens twice a month takes a quarter to say anything. It has to fail recoverably, meaning a person can undo the mistake before a customer is affected. It needs a named owner who will actually watch it. And it should touch few enough systems that the integration work does not dominate the schedule. Inbound inquiry routing satisfies all four in most small companies, which is why it so often goes first, and it also produces the kind of number that makes the second project easy to fund. ### Is the first project mostly about the automation or about learning? Both, and treating it as only the first is the mistake. The first project is the cheapest opportunity you will get to find out three things that no proposal can tell you: how your team reacts to a system making decisions, how many exceptions your process really produces once someone is counting, and whether the vendor tells you about problems before you find them. Those answers change what the second project should be, and sometimes they change whether there is a second project. This is also why we keep the first one small enough to finish in two to four weeks. A six-month first project is a bet placed before any of the information arrives, and it commits you to a supplier and a design at exactly the moment you know the least. ### What if the obvious first candidate touches money? Then split it rather than skipping it, because the part that touches money is rarely the part where the time goes. In an order flow, availability checking, document assembly and dispatch scheduling can be automated while the commitment itself, the price and the terms stay with a person. That gives you the frequency and the time saving without putting an early system in a position to bind the company to something. It is the same rule we apply permanently rather than only at the start: anything binding goes to a human, exceptions route with full context, and every automated decision is logged. Starting with the non-binding portion is not a compromise; it is the design the finished system will have anyway. --- # Five Signs Your Company Should Not Automate Yet URL: https://inite.ai/en/blog/what-size-company-should-not-automate-yet Date: 2026-08-16 Author: Mikhail Savchenko Category: Operations Tags: Operations, Procurement, Automation, Strategy ## Direct Answer Headcount is the wrong test. What decides whether automation pays now is five conditions: enough volume for a fixed build cost to divide into, a process stable enough to survive the payback window, somebody inside who owns the result, data clean enough that automating it does not encode the mess, and a bottleneck that is actually here rather than somewhere else. Failing any one of them is usually a reason to wait a quarter rather than to buy. A twelve-person company can pass all five and a two-hundred-person one can fail three, which is why size predicts so little. ## Key Facts - We put 1-3 workflows into production in 2-4 weeks, and payback typically lands at 3-6 months. - A process that changes materially every month will be rebuilt 3 to 6 times inside that payback window. - The entry point is a 15-minute diagnostic, free, before anyone talks price or scope. - 5 conditions decide readiness, and failing 1 of them is usually enough to wait. ## Size is the wrong question The question comes in a standard shape: are we big enough for this yet? It has no useful answer, because [the arithmetic that decides it does not track headcount](/en/answers/is-my-process-worth-automating). A twelve-person rental business with four hundred bookings a month has more automatable volume than a two-hundred-person consultancy where every engagement is bespoke. Size correlates with a couple of things that matter, and it predicts none of them well enough to use. Five conditions do decide it. Failing one is usually a reason to wait a quarter. ## One: the volume has to divide the build cost Automation economics are a fixed cost spread across throughput. That single sentence explains most projects that disappoint. A process that runs four hundred times a month and saves fifteen minutes each time is a straightforward case. The same process at eighty times a month has the same build cost spread across a fifth of the benefit, and it usually fails honest arithmetic even though the per-run saving is identical. Take the volume from your own systems rather than from an interview, use the low month rather than the good one, and ask what the payback looks like if the volume never grows. The [four questions that break most ROI numbers](/en/blog/roi-math-for-automation-projects) go into the rest of that arithmetic. ## Two: the process has to hold still It needs to be recognizably the same process at the end of the payback window, which for us is three to six months. The test is concrete. Describe the steps as they were six months ago and as they are now. If the difference is in the edge cases, you are fine, because edge cases are what the exception queue is for. If the sequence itself has changed twice, you are automating a design that is still being drawn, and a process that changes materially every month gets rebuilt three to six times before the payback arrives. Waiting for it to settle is much cheaper than rebuilding it, and it is not close. ## Three: someone inside has to own it One named person, whose job makes the workflow theirs after the vendor leaves, with enough authority to change it without a committee. Not the person who signed the contract. Usually not the most senior person in the room. What the owner does is unglamorous: watch the exception queue, notice when the volume shape changes, decide whether a new edge case gets a rule or a human. Automation without an owner degrades quietly, and quietly is the problem. The dashboards keep looking fine while the exceptions pile up and staff invent workarounds around the system. If you cannot name the person before the project starts, fix that first. It costs nothing and it is the single strongest predictor we have. ## Four: the data mess has to be the right kind | Kind of mess | Automate now? | Why | | --- | --- | --- | | Missing fields | Yes | The workflow can ask; gaps become visible when they matter | | Inconsistent formats | Yes | Normalizing is cheap and mechanical | | Stale records | Usually | Automation surfaces staleness faster than people do | | Two systems that disagree | No | Automating a contradiction executes it at speed | The precondition is not clean data. It is a decided answer to which source is authoritative for each field the workflow depends on. Without that decision, automation does not clean the mess, it encodes it, and the disagreement gets executed faster than a person could have caught it. ## Five: the bottleneck has to actually be here The most expensive version of this mistake is a beautifully automated back office attached to a company whose real constraint is that not enough people are asking. Operational automation makes a business faster at fulfilling demand. If demand is the shortfall, it makes you faster at doing less, and the efficiency gain is real and commercially invisible. The gain we measure runs forty to sixty per cent of the automated portion, not of the company, and if the automated portion is not the constraint then the company-level effect is close to nothing. Test it by asking what would happen if the process took half as long. If the answer is that more work would get done and it has revenue attached, this is the right project. If the answer is that everyone would wait more comfortably, the money is better spent on the thing they are waiting for. ## What to do with a "not yet" The right response is rarely to do nothing for a quarter. Measure. Pull the timestamps, count the volume in the low month, name the owner, and decide which system is authoritative for each disputed field. That work is useful whether or not you ever automate, it makes the eventual project shorter, and it is exactly the audit described in [what a process audit must produce](/en/blog/process-audit-before-automation) and demonstrated on a real operation in [the measurement week](/en/blog/rental-case-the-week-before). We say no to projects on these grounds, and it is worth being direct about why that is in our interest as well as yours. A project that fails the arithmetic still fails after we have been paid, and the reference is worth more than the fee. ## FAQ ### Is there a headcount below which automation never makes sense? No, and the fact that people keep looking for one is the reason this question gets answered badly. The arithmetic is dominated by how often a process runs and how long each run takes, and neither of those tracks headcount reliably. A twelve-person equipment rental business handling four hundred bookings a month has more automatable volume than a two-hundred-person consultancy where every engagement is different. What headcount does predict is a second-order thing worth knowing: smaller companies more often lack a person who can own the result after handover, and larger ones more often have processes that three departments disagree about. Both are real obstacles, but they are obstacles about ownership and agreement rather than about size, and they should be tested directly instead of inferred from a number of employees. ### How stable does a process have to be? Stable enough that it will still be recognizably the same process at the end of the payback window, which for us is three to six months. That is a lower bar than it sounds, because most operational processes are far more stable than the people running them believe; what changes weekly is usually the exceptions rather than the main path. The test is concrete: describe the steps as they were six months ago and as they are now, and see whether the difference is in the sequence or only in the edge cases. If the sequence itself has changed twice, you are looking at a process that is still being designed, and automating a design in progress means paying to rebuild it three to six times before it settles. Wait for it to settle. The waiting is cheaper than the rebuilding and it is not close. ### What does 'somebody owns it' actually mean in practice? One named person, inside your company, whose job description makes the workflow theirs after the vendor leaves, and who has enough authority to change it without convening a committee. It is not the person who signed the contract and it is usually not the most senior person involved. What that owner does is unglamorous and decisive: they watch the exception queue, they notice when the volume shape changes, they decide whether a new edge case gets a rule or a human, and they are the one who says something is wrong before the reporting says so. Automation without an owner degrades quietly, because the dashboards keep looking fine while the exceptions pile up and staff invent workarounds. If you cannot name the person before the project starts, that is the thing to fix first, and it costs nothing. ### Our data is messy. Should we clean it first or automate first? It depends entirely on which kind of mess it is, and the distinction is worth ten minutes before deciding. Missing and inconsistent data is usually fine to automate around, because a workflow can be built to ask for what it needs, and in practice automation tends to improve this kind of mess by making the gaps visible at the moment they matter. Contradictory data is the dangerous kind: two systems that disagree about the same fact, with no rule for which one wins. Automating that does not clean it, it encodes it, and the disagreement gets executed at speed instead of being caught by a person who knew the second system was the reliable one. The precondition is not clean data; it is a decided answer to which source is authoritative for each field that the workflow depends on. --- # What the Amazon vs Perplexity Ruling Changed URL: https://inite.ai/en/blog/agentic-browsing-after-the-amazon-ruling Date: 2026-08-15 Author: Mikhail Savchenko Category: Agentic Engineering Tags: Agentic Engineering, Legal, Web Bot Auth, Strategy ## Direct Answer In August 2026 the Ninth Circuit vacated the preliminary injunction that had barred Perplexity's Comet browser from operating on Amazon, holding that Amazon is unlikely to prevail under the Computer Fraud and Abuse Act because on the record before the court it is Amazon's own customers, not Perplexity, who access Amazon's systems. It is the first federal appellate ruling on whether AI agents acting for users may access online platforms. The court limited its holding to that record and declined to state broader principles, so the practical lesson is narrow: the CFAA is a weak instrument against an agent a customer chose to use. ## Key Facts - Amazon filed against Perplexity in November 2025, pleading the federal CFAA and California's CDAFA. - A district court granted Amazon a preliminary injunction in March 2026, reported on 10 March 2026. - The Ninth Circuit stayed that injunction pending appeal and vacated it in August 2026. - It is the 1st federal appellate ruling on whether AI agents acting on a user's behalf may access an online platform. - Web Bot Auth has been supported at AWS WAF since November 2025 and at Cloudflare since early 2026. ## What the court actually held Amazon sued Perplexity in November 2025 over its Comet browser, pleading the federal Computer Fraud and Abuse Act and California's Comprehensive Computer Data Access and Fraud Act. A district court granted a preliminary injunction in March 2026. The Ninth Circuit stayed it pending appeal, and in August 2026 vacated it. The reasoning is the part worth carrying away. On the record before the panel, the systems were being accessed by Amazon's own customers, signed into their own accounts, using software they had chosen. Perplexity was not the one accessing Amazon. On that basis Amazon was unlikely to prevail on a statute written about unauthorised access. It is the first federal appellate ruling on whether AI agents acting for a user may access an online platform, and the panel was careful to say it was deciding that record rather than announcing a doctrine. ## What it did not hold It did not say agents are welcome, and it did not say a site has lost control of its own front door. Contract claims were not what the panel found weak. Terms of service, trademark questions and state-law theories are all untouched. A different record with different facts, particularly one where the agent operates at scale rather than for one signed-in customer, could come out differently. The useful summary is narrow and worth stating without decoration: computer-misuse law is a weak instrument against software a customer chose to run on their own account. ## The distinction the ruling turns on | | Crawler | User's agent | | --- | --- | --- | | Acting for | Its operator | One signed-in customer | | Scale | Many sites, high volume | One session at a time | | Authenticated | Usually not | As the customer | | Data ends up | In the operator's product | In front of the person who asked | | The ruling's reasoning | Does not apply | Applies | Most blocking rules in the wild do not make this distinction. A blanket refusal of automated access catches a customer's own agent alongside the scraper it was aimed at, and those two are commercially opposite: one is a visitor with a wallet, the other is a cost. ## What actually gives you control Three levers, and none of them are statutes. Terms of service are a contract question, and contract was not the theory that failed here. Identity is the second lever: an agent that signs its requests can be recognized and treated deliberately, which is what HTTP Message Signatures and Web Bot Auth exist for, supported at AWS WAF since November 2025 and at Cloudflare since early 2026. Rate and behavior limits are the third, and they are the only ones that keep working regardless of what a visitor claims to be. Together they let you decide per class of visitor on purpose. The alternative is deciding by accident, which is what a blanket rule does. The trade-offs of each choice are set out in [the crawler allowlist post](/en/blog/ai-crawler-allowlist-2026). ## The question for an operator Whether agent-mediated customers are worth having is a commercial decision, not a legal one, and it is answerable with your own data. Find out whether that traffic already reaches you. Then check whether your checkout, booking or inquiry flow can actually be completed by software driving a browser as a signed-in customer, and whether any step depends on a person noticing something on a screen. [What an agent sees on your site](/en/blog/browser-agent-ready-saas) covers the mechanics of that audit. Both failure modes cost money. Silently blocking agent-mediated customers who would have converted turns away business. Letting them arrive and fail halfway through generates support load and a bad impression, which is worse than a clean refusal. ## Where this sits next to the other retreat It is worth reading this alongside what happened to in-chat checkout. OpenAI pulled back from completing purchases inside ChatGPT on 4 March 2026 and confirmed the retirement on 24 March, after roughly five months live and about a dozen Shopify merchants ever going live. Discovery moved into the assistant; the transaction went back to the merchant's own site. The two events point the same way. Buying inside the assistant lost, and the customer's own agent operating the merchant's site just survived its first appellate test. That makes [the merchant's own checkout the surface that matters](/en/analyze), which is the argument made in full in [agentic commerce after Instant Checkout](/en/blog/agentic-commerce-after-instant-checkout). Sources: [Reuters via Yahoo Finance](https://finance.yahoo.com/technology/ai/articles/us-court-overturns-amazon-injunction-135004000.html), [Engadget](https://www.engadget.com/2230471/perplexity-has-successfully-overturned-amazon-injunction-on-its-ai-shopping-bot/), [PYMNTS on the CFAA narrowing](https://www.pymnts.com/news/artificial-intelligence/2026/ninth-circuit-narrows-cfaa-reach-in-perplexity-agentic-commerce-ruling/), [CNBC on the March injunction](https://www.cnbc.com/2026/03/10/amazon-wins-court-order-to-block-perplexitys-ai-shopping-agent.html). ## FAQ ### Does this mean AI agents can now use any site freely? No, and the panel went out of its way to prevent that reading. The holding is that on the factual record presented, Amazon was unlikely to succeed on its Computer Fraud and Abuse Act claim, because the access to Amazon's systems was being performed by Amazon's own customers signed into their own accounts rather than by Perplexity. That is a conclusion about one statute applied to one evidentiary record. The court explicitly declined to announce broader principles about agentic AI or about liability in other legal contexts, which leaves contract claims, terms-of-service claims, trademark questions and state-law theories entirely open. Read as a rule for your own site, it says something much narrower than the headlines: reaching for computer-misuse law against software a customer chose to run is a weak play, and the strength of your position depends on which theory you bring rather than on how you feel about agents. ### What is the practical difference between an agent and a crawler here? Who is acting, and on whose instruction, which turns out to be the hinge the whole ruling swings on. A crawler visits your site for its operator's purposes, usually at scale, usually not signed in, and usually to collect data that will be used somewhere else. An agent in the Comet sense is running for one person, signed into that person's account, doing something that person asked for and could have done by hand more slowly. The reasoning that Amazon's customers rather than Perplexity were the ones accessing the systems only works for the second shape. This matters for how you write your rules: a blanket block on automated access sweeps up your own customers' agents alongside the scrapers you actually meant to stop, and the two have different legal footing and very different commercial consequences. ### So what actually gives a site control over agent traffic? Three things, none of which are computer-misuse statutes. Your terms of service are a contract question rather than a hacking question, and contract theories were not what the panel found weak. Identity is the second: an agent that signs its requests can be recognized and allowed, throttled or refused on purpose, which is the direction the ecosystem is moving with HTTP Message Signatures and Web Bot Auth support at AWS WAF since November 2025 and at Cloudflare since early 2026. Rate and behavior limits are the third, and they are the only ones that work regardless of what anyone claims to be. The combination lets you make a deliberate decision per class of visitor instead of an accidental one, and that decision is a commercial choice about whether agent-mediated customers are worth having. ### Should a small business do anything differently because of this? Most should do one thing, and it is not legal. Find out whether agent-mediated traffic is already reaching you and what it does when it arrives, because you cannot make a decision about a channel you are not measuring. Check whether your checkout, booking or inquiry flow can be completed by software driving a browser as a signed-in customer, and whether anything in it depends on a human noticing something on screen. That is a product question with a commercial answer: if agent-mediated customers convert and you are silently blocking them, you are turning away business, and if they arrive and fail halfway through, you are generating support load. The legal position only becomes relevant after you have decided which of those you want, and for most operators the decision is worth more than the ruling. --- # Where four weeks comes from, and when you will not get it URL: https://inite.ai/en/blog/4-week-vertical-cloning-playbook Date: 2026-08-14 Author: Mikhail Savchenko Category: Operations Tags: Operations, Implementation, Vendor selection, Strategy ## Direct Answer Two to four weeks does not rest on how fast anyone writes code. It rests on how much of the work has nothing to do with your project. Sign-in, roles, permissions, notifications, invoices and message history are identical in any product and were debugged on earlier builds. What gets written fresh is your subject matter, and our own number shows how small a share that leaves: building a second product in an adjacent industry, 7 of 113 business concepts carried over, about 6%. Hence the boundary. The more ordinary your domain, the better four weeks holds; the more unusual it is, the longer it takes. ## Key Facts - Of 113 business concepts across two of our products in adjacent industries, 7 were shared, about 6%. - One process goes into production in 2-4 weeks, and roughly half of that goes to the subject matter. - The first request in the rental case took 4 hours to process; afterwards it took 8 minutes. - Sign-in, roles, permissions, notifications, invoices and message history are 6 subsystems identical in any product, and none is written again. - 4 signs tell you the estimate will run longer: an undocumented process, decisions made case by case, data no program can read, and a requirement nobody else in the industry has. ## Why the number alone tells you nothing [You are quoted two to four weeks](/en/answers/ai-automation-timeline). You will hear the same number from everyone else. On its own it means nothing, because the estimate rests not on the speed at which code gets written but on the volume of work nobody will be doing on your project. That volume is the thing worth asking about. ## What you are not paying for In any system with staff and customers the same parts repeat. Someone signs in under their own account, they have a role, the role permits some things and forbids others, notifications go out to someone, an invoice goes to someone else, and the whole correspondence has to sit in one place and be findable. None of that has anything to do with whether you rent out excavators or book physiotherapy. It was written once, debugged on earlier work, and that is why it takes none of your weeks. There is one exception worth knowing even if you are not the one building. Separation between customers is laid in first or it is never laid at all. Retrofitting it into a system that assumed a single client is the most expensive rework in this business, and it is what turns a four-week project into a three-month one. ## What has to be built either way What is left is the thing you came for: your objects and your process. A rental business has a unit of equipment, a booking and a dispatch. A property agency has a property, a listing and a deal. From the outside it sounds like the same thing. Inside there is almost nothing in common. We measured it on ourselves and the figure was inconvenient. We built a rental platform, then a real-estate platform. Adjacent industries, both about an object handed to somebody for a while or for good. Of 113 business concepts, 7 were shared. About 6%. That number is here so you know what you are buying. "We already have this, we just need to configure it" is a sentence about the other 94% - the ones the vendor does not have, because they are yours. ## When four weeks will not happen | Sign | How it looks at your end | Where the time goes | | --- | --- | --- | | Process undocumented | Three employees describe it three ways | Into weeks of finding out before anything starts | | Decisions made case by case | The rule is not stated even out loud | Into cases that never came up in the demo | | Data unreadable by a program | A mailbox, one person's spreadsheet, memory | Into the integration, or into giving it up | | A requirement nobody else has | "We have always done it this way" | Into building with nothing to build on | None of these makes automation impossible. Each moves work out of weeks of building and into weeks of finding out, and finding out does not transfer: somebody else's earlier project does not know what counts as a request at your company. If even one matches, four weeks stays a fair estimate of the build and a poor estimate of the project. That gap is usually what the argument with the vendor is about a month later. ## What to ask, to check any of this One question separates a verifiable estimate from a nice one: **which of this have you already written, and which part will you write again for us**. "We build everything from scratch for you" means months, because sign-in, roles and permissions have to be written again. "We have it all ready, we just need to configure it" means a template, and you find out where it does not match in week three. The answer worth hearing names both parts separately and shows where the line runs between them. Then ask them to lay that line across your own process. Working through it takes a few hours and, before a quote is signed, is worth more than any presentation: [how to read an automation quote](/en/blog/how-to-read-an-automation-quote) starts from exactly that split. What it looks like on real requests is in the rental write-ups: [where four hours go in a rental request](/en/blog/order-processing-equipment-rental) and [what we measured in the week before](/en/blog/rental-case-the-week-before). If it turns out you are too early to automate, that is a result too: [five signs](/en/blog/what-size-company-should-not-automate-yet) list when waiting is the better call. ## FAQ ### Why does every vendor quote roughly the same number? Because the number is quoted for one process rather than a whole system, and in that sense many of them are being honest. The difference is not the figure, it is what sits behind it. For some, two to four weeks means sign-in, roles, permissions, notifications and invoices are already written and debugged on earlier work, so the whole stretch goes to your subject matter. For others it means a finished template they will try to fit you into, and the estimate holds exactly until the first place your process does not match. For a third group it is a ceiling produced before anyone looked at your actual requests. The question worth asking is not how long, but which of this have you already written and which part will you write again for us. ### What counts as subject matter, and why does it not transfer? It is your objects and your process: what you sell or service, and what happens to it between the first contact and the close. A rental business has a unit of equipment, a booking and a dispatch. A property agency has a property, a listing and a deal. The words look alike and the rules inside are not, and it is the rules that cost time. We measured this on ourselves: we built a rental platform, then a real-estate platform, adjacent industries, and of 113 business concepts 7 were shared. About 6%. Which is why a vendor saying we already have this, we just need to configure it is not describing your project. ### Where do four weeks go if the base is ready? About half goes to the subject matter: writing down your objects, your process and the rules for moving between stages, so the system makes the same decisions your employee would and refuses the ones they are not allowed to make. Another share goes to connecting what you already run - mail, warehouse, accounting, calendar, telephony. Automation with no access to your systems can only repeat what is already published. The rest goes to the edge cases that surface on real requests rather than in a demo. The first week almost always goes not to code but to agreeing what counts as a request and when it is closed. ### How do I know my case will not fit in four weeks? Four signs, and any one of them is enough to plan for more. First: your process is written down nowhere and three employees describe it three different ways. Second: the next step is decided by a person weighing the circumstances, and the rule is not stated even out loud. Third: the data sits somewhere no program can read it - in a mailbox, in someone's head, in another department. Fourth: you have a requirement nobody else in the industry has, and it is not up for discussion. None of these makes automation impossible, but each moves work out of weeks of building and into weeks of finding out, and finding out cannot be inherited from an earlier project. --- # Agency or Freelancer for an AI Automation Build URL: https://inite.ai/en/blog/ai-automation-agency-vs-freelancer Date: 2026-08-13 Author: Anton Fenix Category: Comparison Tags: Comparison, Operations, Procurement, Automation ## Direct Answer For a single well-understood workflow with a named owner inside your company, a freelancer is usually the right call and often the cheaper one by a wide margin. An agency earns its premium on three specific things: work that spans several systems and therefore several skills, a delivery that has to survive one person being unavailable, and an obligation to keep running after handover. The comparison people make first is the day rate, and the day rate is the least decisive number in it. What decides the outcome is who is responsible in month four, when the integration changes underneath the workflow and the person who built it has moved on. ## Key Facts - A single-workflow build is where a freelancer competes hardest, and it is also the shape of roughly 1 of the 1-3 workflows we put into production per engagement. - Our own delivery window is 2-4 weeks per engagement, which is short enough that continuity risk concentrates after handover rather than during the build. - Payback on automation work typically lands at 3-6 months, which is after the point at which a freelance engagement has usually ended. - Every build is preceded by a written ROI estimate at conservative inputs; if it does not clear positive, the engagement ends at the diagnostic. ## The day rate is the wrong first comparison Most of these decisions start with two numbers side by side, one considerably smaller than the other, and end with a discussion about whether the larger one is justified. The day rate is the least decisive number in the comparison. Both parties can usually build the workflow. What separates them is what the thing does in month four, when a supplier changes a form, a channel changes an API, or the volume doubles and an assumption that held at fifty orders a day stops holding at a hundred. ## Where a freelancer wins outright One process. Few systems, all documented. Someone inside your company who will own the result and can change it. Under those three conditions you are buying construction, not a relationship. The specification can be written down completely before anyone starts, the coordination an agency carries is doing no work, and the price difference is large and real. A good freelancer will often be faster as well, because there is no internal handover in a team of one. This is a more common shape than vendors like to admit. If your automation is a single well-understood workflow, the honest advice is to get quotes from individuals. ## Where the premium is actually earned | What you need | Freelancer | Agency | | --- | --- | --- | | One workflow, documented systems | Strong fit | Overqualified | | Four systems, four skill sets | Depends on the individual | Structural fit | | Fixed date tied to a season | Single point of failure | Absorbs absence | | Someone accountable in month four | Personal commitment | Contractual obligation | | Willingness to say "do not build this" | Varies | Varies | The last row is deliberately unresolved, because company size predicts nothing about it. An agency whose diagnostic always concludes that its own product is needed is worse than a freelancer who tells you the truth, and the reverse is equally common. Continuity is the row that people underweight and then regret. A build is 2-4 weeks; the system lives for years. The question is not who writes it, it is who answers the phone when it stops working. ## The cost comparison that is actually useful Build cost favors the freelancer, usually by a lot. Twelve months of ownership is much closer. Every automation carries running costs regardless of who built it. Model and infrastructure spend per decision. The time of whoever handles escalated exceptions, which is a real cost by design rather than a defect. And maintenance when the world around the workflow moves, which it does. Ask both parties for a monthly running figure in writing, add twelve of them to the build price, and compare those totals. That single change makes most of these decisions obvious in either direction, and the [four questions that break most ROI numbers](/en/blog/roi-math-for-automation-projects) apply to both quotes equally. ## The arrangement worth considering Split the work. Have the measurement and design done by whoever you trust to tell you the project is not worth doing. Have the construction done by whoever is cheapest for that shape of work. The halves have different failure modes. Bad construction is cheap and obvious within days. A bad design is expensive and invisible for months. Splitting them also gives you a specification that several builders can quote against, which is the only way to make the price comparison mean anything. The reverse arrangement is the one to avoid: hiring a builder first and then asking them what should be built puts the scoping decision with the party whose income scales with the scope. That is the same structural problem described in [what a process audit must actually produce](/en/blog/process-audit-before-automation). ## What neither of them should get away with Two things, and they apply identically to individuals and firms. Neither should quote before measuring. A price produced from a conversation is a guess about a process nobody has counted, and the [measurement week](/en/blog/rental-case-the-week-before) that precedes an honest quote is cheap enough that skipping it is a choice rather than a constraint. And neither should leave the decision about what stays human implicit. Anything binding, anything unusual, and anything the system is unsure about belongs with a person, and that boundary should be in the proposal rather than discovered later. For how we structure our own engagements against these criteria, the [comparison page](/en/compare) sets out what we do and do not take on. ## FAQ ### When is a freelancer clearly the better choice? When the workflow is one process, the systems it touches are few and well documented, and somebody inside your company will own the result and is technical enough to change it. Under those conditions you are buying construction rather than a relationship, the specification can be written down completely before work starts, and the price difference is large and real. A good freelancer will beat an agency on cost and often on speed for exactly this shape of work, because none of the coordination an agency carries is doing anything useful here. The failure mode to watch is scope discovery: if the process turns out to touch four systems nobody mentioned, a solo engagement can stall in a way a team absorbs. That is a reason to spend a week measuring before contracting, not a reason to hire a bigger vendor by default. ### What exactly is the agency premium paying for? Three things, and it is worth checking that you actually need each of them before paying for all three. Breadth first: a workflow that spans a CRM, an accounting system, a messaging channel and a document store needs several kinds of knowledge, and one person good at all four is rarer and more expensive than a team that has them separately. Continuity second: an agency can lose a person mid-build without losing the build, which matters more the more the timeline is tied to a season or a deadline. Obligation third, and this is the one people underweight: someone contractually still there in month four when a supplier changes a form or a channel changes an API. A freelancer can offer all three and some do, but they are offering them as a personal commitment rather than as a structure, and personal commitments end when circumstances change. ### How do the two compare on cost once running costs are included? The build cost usually favors the freelancer clearly, and the total cost over the first year is much closer than the day rates suggest. Automation carries running costs regardless of who built it: model and infrastructure per decision, the time of the person handling escalated exceptions, and maintenance when the surrounding systems move. Those costs exist in both cases; the difference is who absorbs the maintenance and at what notice. A freelancer typically prices maintenance as new work at a new rate and subject to availability, which is fine when the system is stable and expensive when it is not. Compare the two on twelve months of ownership rather than on the build, and ask both parties for a monthly running figure in writing before signing anything. ### Can you mix the two? Yes, and it is a sensible pattern that gets used less than it should. Have the measurement and the design done by whoever you trust to tell you the project is not worth doing, then have the construction done by whoever is cheapest for that shape of work. The two halves have different failure modes: a bad design is expensive and invisible for months, while bad construction is cheap and obvious within days. Splitting them also gives you something to compare, because a design written down properly can be quoted by several builders. The one arrangement to avoid is the reverse, where a builder is hired first and then asked to decide what should be built, since that puts the scoping decision in the hands of the party paid by the size of the scope. --- # The Week Before: What We Measured at a Rental Firm URL: https://inite.ai/en/blog/rental-case-the-week-before Date: 2026-08-12 Author: Mikhail Savchenko Category: Case Study Tags: Case Study, Operations, Equipment Rental, Process Audit ## Direct Answer Before automating anything at an equipment rental operator we measured for a week, and the measurement changed the project. Average booking-to-dispatch time turned out to be the least useful number we collected, because the average hid the tail and the tail was where refused bookings lived. The counts that decided the build were how many inquiries arrived outside working hours, how many were never answered at all, and how often a machine standing in the yard was reported as unavailable. The rebuild that followed took 3 weeks and moved booking-to-dispatch from 4 hours to 3 minutes. ## Key Facts - The rebuild that followed the measurement week took 3 weeks and moved booking-to-dispatch from 4 hours to 3 minutes. - Peak-season capacity rose 2.5x afterwards, with no additions to the team. - Of the 4 numbers we expected to drive the business case, 2 turned out not to matter. - We put 1-3 workflows into production in 2-4 weeks, and this project sat at the short end of that range. - Every build is preceded by a written ROI estimate at conservative inputs; if it does not clear positive, the engagement ends at the diagnostic. ## Nobody knows their own numbers The rental operator we worked with could describe the problem precisely. Bookings were handled by hand, availability lived in spreadsheets, contracts were assembled one at a time, and drivers were arranged by phone. Peak season meant lost bookings and double-bookings. Every part of that description turned out to be true. None of it was a measurement. So the first week of the project produced no software. It produced a count, taken from the systems the business already had, of how big each part of the problem actually was. That week is the reason the build took three and not six, because it removed two things from scope that everyone had assumed were central. ## What we pulled Two timestamps per booking, across a full peak month: when the inquiry arrived, and when dispatch was confirmed. Then three counts that are not timestamps. - Inquiries that arrived outside working hours - Inquiries that were never answered at all - Occasions when a machine physically present in the yard was reported unavailable None of this needed new instrumentation. It needed somebody to export what was already recorded and count it without flinching. ## The average was the least useful number Booking-to-dispatch averaged around four hours, which is the number that ended up in every later description of the project including ours. It was also the number that mattered least during scoping. | What we looked at | What it told us | | --- | --- | | Average time to dispatch | The problem exists | | Distribution of that time | Where the money is | | Share arriving out of hours | Why the tail is long | | Inquiries never answered | What was being lost silently | | Yard-present but unavailable | Whether the data model was at fault | Most bookings moved at a reasonable pace. A minority waited far longer than the average suggested, and that minority was where customers stopped waiting and called a competitor. An improvement to the average would have read well in a report and changed nothing commercially, because the customers who left were never in the middle of the distribution. This generalises past rental. The average is the number most likely to survive into a proposal and least likely to identify the constraint. ## Two things we expected to matter and dropped Weekly booking volume was the first. The annual figure flattened a season that earns most of the year in a few weeks, and scoping against the annual number would have sized the system for a load it never sees when it matters. The second was contract preparation time. It was real, it was tedious, and it was not the constraint. Document assembly is genuinely slow work, and automating it saves exactly the minutes it takes, which is a small share of the four hours. It stayed in the build because it is cheap once the rest exists. It stopped being the reason for the build. Removing both from the critical path is most of why the project fit into 3 weeks. ## The number that decided the shape The count of machines standing in the yard and reported unavailable was the one that changed the design. Had that count been near zero, the honest conclusion would have been that staff were overloaded and the fix was capacity or routing. It was not near zero. The availability data itself was wrong often enough to explain both the double-bookings and a share of the refusals, which pointed at the model rather than at the people. That is what made the build mostly a data-modeling job with a language model at the front door rather than the reverse, and [where the four hours go](/en/blog/order-processing-equipment-rental) sets out the resulting split between rules and model in detail. ## What happened next The rebuild took 3 weeks. Booking-to-dispatch went from 4 hours to 3 minutes, double-bookings were eliminated completely, and peak-season capacity rose 2.5x with the team that was already there. The [case page](/en/cases/equipment-rental-automation) carries the full set. The capacity figure is the one that pays for the project, and it traces directly back to the measurement week. The refused bookings were countable before anything was built, which is why the business case did not depend on believing a forecast. ## If you want to run this week yourself You do not need us to do it, and there is a reasonable argument that you should do it before talking to any vendor at all. Pull the two timestamps, count the three things, and look at the distribution rather than the average. The measurement is the first stage of [the audit we run before agreeing to build anything](/en/blog/process-audit-before-automation), and the numbers it produces are the ones that make [the four questions that break most ROI claims](/en/blog/roi-math-for-automation-projects) answerable rather than rhetorical. If the tail is thin and nothing was being turned away, the answer is that this is not a project worth doing. We would rather tell you that in week one than in month three. ## FAQ ### Why measure for a week instead of just asking the team? Because people are accurate about what a task feels like and unreliable about how often it happens and how long it waits. Ask a coordinator how long a booking takes and you get the duration of their own part of it, which is genuinely short, plus an impression of the waiting, which is genuinely long but not something anyone tracks. Neither answer is dishonest and neither is usable for a business case. Timestamps from the systems are different in kind: they show the distribution rather than the impression, they cover the nights and weekends nobody remembers, and they include the inquiries that were never answered, which by definition nobody can recall. The week is also short enough to be affordable and long enough to catch the shape. We would rather spend a week finding out that a project is not worth doing than spend three weeks building the thing that was not the constraint. ### What exactly did you pull, in practical terms? Two timestamps per booking across a full peak month, taken from whatever systems already held them: when the inquiry arrived, and when dispatch was confirmed. Then three counts that are not timestamps at all. How many inquiries arrived outside working hours, because those are the ones that sit until morning and are invisible in any average that mixes them with the rest. How many inquiries received no response ever, which is the number nobody wants to look at and the one most likely to contain revenue. And how many times a machine physically present in the yard was reported as unavailable, which is the signal that the availability model rather than the staff is the problem. None of that requires new instrumentation. It requires somebody to export what is already recorded and count it honestly. ### Which of the numbers turned out not to matter? The average booking-to-dispatch time and the count of bookings per week. Both were the numbers everyone expected the case to rest on, and both turned out to be nearly useless on their own. The average hid the distribution: most bookings moved reasonably fast and a minority waited a very long time, and it was that minority where customers gave up and went elsewhere. Improving the average would have looked good in a report and changed nothing commercially. Volume was similarly misleading because the yearly figure flattened a season that does most of the earning in a few weeks. What decided the build was the shape of the tail during peak and the count of inquiries that were never answered. This is not specific to rental: the average is the number most likely to survive into a proposal and least likely to identify the constraint. ### What if the measurement says the project is not worth doing? Then we say so and do not build it, and that has to be a real possibility for the measurement to mean anything. The rule we work by is that if the audit cannot show a return, there is no build, and it exists because the alternative is a vendor whose diagnostic always concludes that its own product is required. Concretely for a rental operation, the case falls apart when the peak-month tail is thin and nothing was being turned away. In that situation the four hours between booking and dispatch are an irritation rather than a cost, the saved hours would not become capacity anyone sells, and the honest recommendation is to spend the money on something else. Turning that project down costs us a contract and keeps the diagnostic worth trusting, which is the only reason anyone lets us measure their operation in the first place. --- # Where the Four Hours Go in a Rental Booking URL: https://inite.ai/en/blog/order-processing-equipment-rental Date: 2026-08-11 Author: Mikhail Savchenko Category: Automation Tags: Automation, Operations, Equipment Rental, Order Processing ## Direct Answer Order processing in equipment rental is slow because the work is short and the waiting is long. A booking sits between a coordinator opening a spreadsheet, someone confirming the machine actually came back, an agreement being drafted, a signature being chased, and a driver being called. Each step takes minutes; the gaps between them take hours. Automating the steps saves minutes, and automating the handoffs is what moved one rental operator from four hours to three minutes. Most of that fix was deterministic rules rather than a language model, because an availability check has to answer identically every time. ## Key Facts - At one equipment rental operator, booking-to-dispatch fell from 4 hours to 3 minutes and double-bookings stopped completely. - The same operator absorbed 2.5x its peak-season volume with the staff it already had. - That build took 3 weeks, inside the usual window of 2-4 weeks for 1-3 workflows in production. - Of the 7 asset states in the availability table below, 4 leave a machine physically present in the yard and still not rentable. - A first delivery is 1-3 workflows in production in 2-4 weeks, handed to the client's own team with documentation and monitoring. ## The four hours were never one task Ask a rental coordinator how long it takes to turn a booking into a dispatched machine and the answer is usually a shrug and a number. Four hours, most days. Longer in July. Break the four hours apart and almost none of it is work. An inquiry arrives, in one of four or five places. Someone opens the availability sheet. Someone else has to confirm that the machine which was due back yesterday actually came back. An agreement gets drafted from the last similar one. A signature is chased. A driver is called, and the driver is on a job. Each of those steps takes a few minutes of somebody's attention. The four hours are the gaps between them, waiting for the next person to be free. This distinction decides what an automation project is worth. Automating the steps saves the minutes, and the minutes were never the problem. Automating the handoffs is what closes the gap, and that is the change that took one operator from four hours to three minutes. ## Two people booking the same machine is a concurrency problem Double-booking gets treated as carelessness, which is why the usual remedy is a reminder to be careful. It does not work, and the reason it does not work is worth understanding before buying anything. A shared spreadsheet has no locking. Two coordinators open it at 10:02. Both see the 3-tonne excavator free on Thursday. Both promise it, one on the phone and one by email. Each was individually correct at the moment they looked, and neither had any way to know the other was halfway through a promise. The error was created by the tool, not by the people. Fixing it needs two things: one record that is the single writer of truth about what is reserved, and a check that runs at the moment of commitment rather than only at the moment of inquiry. The window between looking and promising is where the conflicts live. ## Availability is not a yes-or-no field The second reason bookings collide is that most systems store availability as one boolean, and a rental asset has more states than that. | State on Thursday morning | In the yard | Rentable Thursday | | --- | --- | --- | | On rent, due back Wednesday | No | Yes, if the return holds | | On rent, job overrunning | No | No | | In transit back from site | No | Depends on the distance | | Returned, awaiting inspection | Yes | No | | In maintenance | Yes | No | | Reserved but unconfirmed | Yes | No | | Present and free | Yes | Yes | Four of those seven rows describe a machine standing in the yard that cannot go out. A check that asks only whether the asset is on a live rental contract answers "available" for all four. That is the bulk of the double-bookings, and it explains why the fix is mostly a data-modeling exercise rather than an AI one. The availability question has to be asked against every state an asset can occupy, including the ones where it is visible from the office window. ## Most of the fix is not a language model The interesting part of this project, commercially, is how little of it is AI in the sense the word is usually sold. | Step | Handled by | Why | | --- | --- | --- | | Availability and conflict check | Rules | Must answer identically every time, and be auditable | | Pricing from the rate card | Rules | A price is a commitment, not an inference | | Agreement and confirmation documents | Templates | The approved wording has to stay the approved wording | | Dispatch sequence and notifications | Rules | Deterministic ordering, no judgment involved | | Reading a free-text inquiry | Model | Unstructured input a human would otherwise retype | | Classifying an unusual request | Model, then a person | Judgment, so it routes rather than decides | Deterministic steps are handled by rules precisely because they must behave identically every time. A probabilistic availability check is a defect wearing a fashionable name. The model earns its place at the front door, where an inquiry arrives at eleven at night as a message naming a machine in the customer's own words with the dates written in a form no field expects. Turning that into a structured request is hard for a rule and easy for a model. Keeping the boundary sharp is also what makes the system explainable later, because you can always say which half produced a given answer. ## What stays with a person Four things route to a human by design, and the reasoning behind each is set out in [our rules for keeping a person in the loop](/en/blog/safe-ai-framework-human-in-loop). Anything binding goes to a person: a discount outside the rate card, terms that differ from the standard agreement, anything that commits the company. Exceptions arrive with the full context attached rather than as a bare alert, because an alert without context just moves the work. Documents are assembled from approved templates rather than composed. And every automated decision is logged, so the question of why the system did what it did on a particular Thursday has an answer. A booking that would leave no buffer before the next job is the case worth calling out, because it looks automatable and is not. Whether a two-hour turnaround is acceptable depends on the customer, the site, and how the last job with them went. ## What it was worth The numbers from the deployment, in full: booking processing went from 4 hours to 3 minutes, double-bookings were eliminated completely, staff were freed to handle customers rather than the calendar, and peak-season capacity rose 2.5x. It shipped in 3 weeks. The details are on the [equipment rental case page](/en/cases/equipment-rental-automation). The capacity figure is the commercially interesting one, and it is worth being precise about where it comes from. The hours themselves did not become money. They became money because peak-season bookings were previously being refused for lack of turnaround, and refused bookings turned into accepted ones. That is the first of the two routes by which saved hours reach the accounts, and the [four questions that break most ROI numbers](/en/blog/roi-math-for-automation-projects) are worth running over any figure of this shape, ours included. ## Where to start Measure the gap before pricing the fix, and measure it from the system rather than from an interview. Pull two timestamps for every booking across a full peak month: when the inquiry arrived, and when dispatch was confirmed. Look at the distribution rather than the average, because the average hides the tail and the tail is where refused bookings live. Then count how many inquiries were answered outside working hours, and how many were never answered at all. That measurement is the first stage of [the audit we run before agreeing to build anything](/en/blog/process-audit-before-automation), and it is also the thing that tells you whether the case is real. If the peak-month tail is thin and nothing was being turned away, the honest answer is that the four hours are an irritation rather than a cost, and the money is better spent elsewhere. For what the deployment covers in a rental operation specifically, see [AI automation for equipment rental](/en/industries/equipment-rental) and the [order processing workflow](/en/automation/order-processing). ## FAQ ### How are double-bookings actually prevented? By moving the availability check from a person's memory to a rule, and by moving the moment of the check from when someone looks to when someone commits. A shared spreadsheet has no locking, so two coordinators can read it at the same minute, both see an excavator free on Thursday, and both promise it. Neither of them was careless; each was individually correct at the moment they looked, and the tool had no way to tell either one that the other was mid-promise. The fix has two halves and both are required. One record becomes the single writer of truth for what is reserved, so there is no second copy to disagree with it. And the check runs again at the instant the booking is confirmed rather than only when the inquiry was answered, which is the window the conflicts were living in. This part of the system is deliberately not a language model. It has to return the same answer for the same question every single time, and that is what rules are for. ### Do we have to replace our rental software for this? No, and in most rental operations that would be the expensive way to fix a cheap problem. Nearly every operator above a handful of machines already runs something for contracts and stock. The hours are rarely lost inside that system; they are lost in the manual relay around it, between the system, the shared inbox, the phone, and whoever happens to be standing in the yard. So the automation usually sits on top of what exists and integrates with it, and the work is in the joins rather than in a migration. This matters commercially as well as technically: a replacement project is measured in months and carries the risk of losing history, while closing the relay is measured in weeks and leaves the record where it already is. We deploy 1-3 workflows into production in 2-4 weeks precisely because the scope stays at the joins. ### What should be a rule, and what should be a model? The dividing line is whether the step is allowed to exercise judgment. Availability, conflict detection, pricing from a rate card, document assembly from approved templates, and the sequence of dispatch steps are all deterministic. They must behave identically on identical input, they have to be auditable afterwards, and a probabilistic answer is a defect rather than a feature. Those are rules. The language model earns its place where the input is unstructured and a human would otherwise have to read and retype: an inquiry arriving as free text in a messenger at eleven at night, naming a machine in the customer's own words, with dates written in a form no field expects. Turning that into a structured request is genuinely hard for a rule and genuinely easy for a model. Keeping the boundary sharp is what makes the system explainable when something goes wrong, because you can always say which half produced the answer. ### We are seasonal. Is a three-week build worth it? Seasonality is the argument for doing it rather than against. Peak season is when a rental business earns most of its year, and manual dispatch is usually what caps how much of that peak it can accept. When the queue grows faster than the coordinators can work through it, the constraint stops being the fleet and starts being the relay, so bookings get refused while machines sit available. That is the situation the 2.5x capacity figure came out of: the operator was already turning peak work away, and closing the booking-to-dispatch gap turned refused bookings into accepted ones. The build takes 2-4 weeks, so a project started in the off-season is running before the season opens. Starting it during the peak is the version that does not pay, because the people who have to answer questions during the build are the same people the peak is already consuming. ### What still needs a person? Anything binding, anything unusual, and anything the system is not confident about. A discount outside the rate card, terms that differ from the standard agreement, a customer with an open damage dispute, and a booking that would leave no buffer before the next job all go to a human, with the full context attached rather than as a bare alert. That is a deliberate design rule and not a limitation we are apologizing for: automation that quietly commits a company to a price or a contract is a liability, and one bad commitment costs more than the hours saved by making it automatically. Documents are assembled from approved templates rather than composed freely, so the wording a lawyer signed off stays the wording that goes out. And every automated decision is logged and auditable, which is what makes it possible to answer the question that always eventually gets asked, which is why the system did what it did on a particular Thursday. --- # The AI Checkout Got Cancelled. The AI Customer Did Not. URL: https://inite.ai/en/blog/agentic-commerce-after-instant-checkout Date: 2026-08-10 Author: Mikhail Savchenko Category: Strategy Tags: Agentic Commerce, Strategy, Operations, E-commerce ## Direct Answer In-chat checkout inside ChatGPT lasted about five months. It launched on 29 September 2025 with Etsy, added a handful of Shopify brands, and was pulled back on 4 March 2026, with OpenAI confirming the change on 24 March. Three ordinary operational problems killed it: sales tax, fraud prevention, and keeping inventory accurate in real time across merchants. What survived is the part that matters for a shop owner. The Agentic Commerce Protocol is still published under Apache 2.0, AI assistants still send buyers to merchants, and the purchase now completes on your own site. That moves the work back to where it already was, which is your product data, your stock accuracy, and your checkout. ## Key Facts - In-chat Instant Checkout ran roughly 5 months: live on 29 September 2025, pulled back on 4 March 2026, confirmed retired on 24 March 2026. - Only about 12 Shopify merchants ever went live with in-chat checkout. - OpenAI's 4% checkout fee, announced for late January 2026, took the combined rate with card processing to roughly 9.2%. - 3 operational gaps ended it: sales tax collection, fraud prevention, and real-time inventory sync across merchants. - The Agentic Commerce Protocol was published under Apache 2.0 on 29 September 2025 and is still maintained by OpenAI and Stripe. ## What actually happened For most of 2026 the advice to online sellers was that buying was about to move inside the chat window and you should prepare for it. That version of the future ran for about five months and then stopped. | Date | What happened | | --- | --- | | 29 Sep 2025 | Instant Checkout launches with Etsy; the Agentic Commerce Protocol is published under Apache 2.0 | | Late Jan 2026 | A 4% checkout fee takes effect, roughly 9.2% combined with card processing | | 4 Mar 2026 | OpenAI pulls back from operating checkout inside ChatGPT | | 24 Mar 2026 | The retreat is confirmed in the updated shopping announcement | Only around a dozen Shopify merchants ever went live. The protocol stayed, the buying behavior stayed, and the payment step went back to the merchant's own site. If you spent this spring being told to get ready for in-chat checkout, that is why the advice went quiet. ## Why it broke The reasons are worth reading closely, because they are the same reasons most operational software is harder than its demo. **Sales tax.** It is calculated per jurisdiction and per product category, and whoever takes the payment carries the obligation. Doing that correctly on behalf of thousands of merchants at once is a genuinely difficult problem, and getting it wrong is expensive in a way that surfaces months later. **Fraud.** The signals that catch a bad order live with the merchant: order history, address patterns, what normal looks like for this catalog. An intermediary sitting in front of many merchants has less of that context and still has to eat the chargebacks. **Inventory truth.** A purchase needs stock to be correct at the moment of purchase, not at the last sync. Selling something you do not have is worse than not selling it, and syncing live availability across many catalogs at scale is exactly the kind of unglamorous plumbing that decides whether a system works. None of these is a verdict on AI shopping. They are a reminder that [the boring machinery around a process is usually the load-bearing part](/en/blog/business-process-automation). ## What survived Three things, and they are the ones that matter to a shop owner. The protocol is still there. ACP remains published under Apache 2.0 and maintained by OpenAI and Stripe, which means the integration you might build is against an open specification rather than one company's product decision. The buyer is still there. People ask an assistant what to buy, compare options in the conversation, and arrive on a merchant site with the decision largely made. That did not depend on where the card details were entered. And the work is still yours. Discovery happens in the assistant, checkout happens on your site. Which is to say the responsibility landed back exactly where it was before anyone promised to take it away. ## What this actually asks of you The assistant never sees your design. [It reads your facts](/en/analyze), and anything it cannot state confidently it leaves out of the comparison entirely, without telling you. That makes the job concrete: - **Facts in machine-readable form.** Price, availability, variants, dimensions, materials, delivery window, returns policy. On the page, in structured data, not only in the visual layout and definitely not only in a photographed spec sheet. - **Stock status that is true right now.** A recommendation that lands on an out-of-stock page costs the sale and the next comparison too. - **Boring pre-purchase answers in plain text.** Sizing, compatibility, what is in the box, how returns work. Assistants read pages, not chat widgets and not PDFs. - **One canonical page per product.** If four URLs describe the same thing, an assistant will pick one, and it may not be the one you would have picked. This is the same discipline that makes a site legible to [any automated visitor rather than a human one](/en/blog/browser-agent-ready-saas), and none of it is wasted if the payment step moves again. ## The honest read on the economics The 4% fee is a useful number to keep, even though the flow it applied to is gone, because it is what an intermediary thought the introduction was worth. Against a marketplace taking 25-30%, that is cheap. Against selling from your own site at roughly 4% all in, it is more than double. Neither comparison is the real question. The real question is whether the sale would have happened without the introduction, and whether the customer becomes yours afterwards or stays the intermediary's. An introduction you convert into a repeat customer is worth paying a lot for. An introduction where somebody else keeps the relationship is a rented channel, and rent is not famous for going down. That is the same arithmetic you would apply to any other channel, which is to say [the payback question does not change because the channel is new](/en/blog/roi-math-for-automation-projects). ## What to do this quarter Nothing dramatic, which is usually the correct answer after a hype cycle deflates. Fix the product data, because it pays off in ordinary search too. Make stock accuracy real rather than nominal. Keep watching where your traffic says it came from. And treat any vendor pitching an urgent agentic commerce integration with the skepticism the last twelve months earned: the standard is open, the buyers are real, and the checkout is on your own site, where it was to begin with. Sources: [OpenAI on Instant Checkout and ACP](https://openai.com/index/buy-it-in-chatgpt/), [Stripe on the open standard](https://stripe.com/blog/developing-an-open-standard-for-agentic-commerce), [the ACP specification](https://www.agenticcommerce.dev/), and [Forbes on the March 2026 retreat](https://www.forbes.com/sites/jasongoldberg/2026/03/10/why-openais-checkout-retreat-spells-trouble-for-its-commerce-strategy/). ## FAQ ### Should I still do anything about agentic commerce, given the checkout was switched off? Yes, but the work is different from what most people were told to do. The part that got cancelled was the payment step happening inside the chat window. The part that did not get cancelled is people asking an assistant what to buy and then arriving on your site with the decision mostly made. That behavior is now the normal front end of a purchase for a growing share of shoppers, and it rewards different things than search did. An assistant comparing three products reads structured facts: price, availability, dimensions, materials, returns policy, delivery window. If your product page states those clearly in text and in structured data, you get compared accurately. If they live only in a photo of a spec sheet or in a PDF, you get skipped without ever knowing. That is the work, and it is the same work whether or not the payment ever moves back into the chat. ### Why did in-chat checkout fail if the demand was there? For three reasons that have nothing to do with AI and everything to do with running a shop. Sales tax is calculated per jurisdiction and per product category, and the party taking the payment carries the obligation, which is a genuinely hard problem to solve on behalf of thousands of merchants in one flow. Fraud prevention depends on signals the merchant has and the intermediary does not, and chargebacks land somewhere real. And inventory has to be correct at the second of purchase, not at the last sync, because selling something you do not have is worse than not selling it. Only around a dozen Shopify merchants ever went live, which tells you the integration cost was high relative to the volume it returned. None of that is a verdict on AI shopping. It is a reminder that the unglamorous parts of commerce are load-bearing. ### What does the 4% fee tell me about the economics? It tells you what an intermediary thinks the introduction is worth, and it is worth comparing against your existing channels before assuming it is expensive or cheap. OpenAI announced a 4% checkout fee for late January 2026, which stacked on top of card processing to reach roughly 9.2% all in. Against a marketplace taking 25-30% that is inexpensive. Against selling directly from your own site at around 4% all in, it is more than double. The right comparison is not the headline percentage but the incremental margin on a sale you would not otherwise have made, and whether the customer becomes yours afterwards or belongs to the intermediary. Introductions you can convert into a repeat customer are worth paying for. Introductions where the intermediary keeps the relationship are a rented channel, and rent goes up. ### How do I make my catalog legible to an AI assistant? Start by assuming the assistant never sees your design, only your facts, and that anything it cannot state confidently it will leave out of a comparison. In practice that means four things. Put price, availability, variant options, dimensions, materials, delivery window and returns policy in machine-readable structured data on every product page, not only in the visual layout. Keep stock status truthful and current, because a recommendation that leads to an out-of-stock page costs you the visit and the credibility. Write the answers to the boring pre-purchase questions in plain text on the page rather than in a chat widget or a PDF, since that is where an assistant reads. And keep one canonical page per product so an assistant does not have to guess which of your four similar URLs is the real one. --- # What Your AI Is Not Allowed to Decide URL: https://inite.ai/en/blog/safe-ai-framework-human-in-loop Date: 2026-08-03 Author: Olga Fedotova Category: Operations Tags: Safe AI, Automation, Operations, Governance ## Direct Answer Safe AI in a working deployment comes down to a few rules about where a person still has to sit. In ours there are four. Nothing that binds the company - a price, a commitment, a contract term - goes out without a person approving it. Anything the system is unsure about reaches a person with the full conversation attached, not as a bare ticket. The system assembles documents from approved templates and validated data instead of composing clauses. Every automatic decision is written down in a form somebody can audit later. Where the approval line sits is decided per process and per jurisdiction during the diagnostic: a clinic in one country and a brokerage in another need different answers. ## Key Facts - In a clinic deployment, intake went from 45 minutes to 8 while every clinical judgment stayed with a clinician. - A brokerage cut first response from 6 hours to 8 minutes and document preparation from 2 days to 20 minutes, with a person still approving what went out. - A first delivery is 1-3 workflows in production in 2-4 weeks, handed to the client's own team with documentation and monitoring. - We put 1-3 workflows into production in 2-4 weeks, and the escalation path is designed in the same 2-4 weeks. - No-shows fell 40% at the clinic, on reminders that require no approval at all. ## The question underneath "is it safe?" Every operator asks some version of it before signing, and it almost never means what the vendor answers. The vendor hears "will the model hallucinate" and starts talking about accuracy rates. The operator means something narrower and much more practical: what is this thing allowed to decide without me? That is a design question with a written answer, and it takes about an hour to settle per process. Here is ours, as four rules rather than a set of principles, because a principle cannot be checked on a Tuesday afternoon and a rule can. ## Rule one: nothing binding goes out without a person A price. A delivery date. A term. A document that a court would read as a commitment. All of it stops at a person. This is the rule that survives contact with lawyers, and it is also the one that costs the least, because the binding step is a small fraction of any process. In a brokerage deployment, the AI answered first, qualified the inquiry, assembled the document package and routed it to the right agent. First response fell from 6 hours to 8 minutes and document preparation from 2 days to 20 minutes. What left the building still had a human name attached to the decision. The six hours were queue time. Nobody was thinking for six hours; the inquiry was sitting in an inbox. Automation is very good at removing waiting, retyping and looking-up, and quite bad at carrying liability. Splitting those two is most of the design. ## Rule two: exceptions arrive with the context attached Any system that escalates uncertain cases to a human [is spending that human's time by design](/en/protocol). The variable is what shape the escalation arrives in. An escalation that says "needs review" makes the approver rebuild the whole situation before deciding, and that reconstruction is usually longer than the decision. An escalation that arrives with the full conversation, the data the system used, and what it was uncertain about turns a five-minute reconstruction into a thirty-second judgment. This matters financially, not just ergonomically. Escalation rate times approver time is a monthly cost that runs forever, and it belongs in the [arithmetic before anyone signs](/en/blog/roi-math-for-automation-projects) rather than being discovered in month three. ## Rule three: the system assembles, it does not compose For anything contractual, the AI works from templates you have approved and data that has been validated. It fills, it selects, it arranges. It does not write new clauses. This is a narrower promise than "the AI drafts your contracts" and it is the reason the legal review is short. A reviewer checking whether the right approved template was used with the right validated data is doing a fast, bounded job. A reviewer checking whether a generated paragraph creates an unintended obligation is doing a slow, unbounded one. The same logic covers regulated content generally: administrative work runs unattended, professional judgment does not. In a [clinic deployment](/en/industries/clinics), patient intake went from 45 minutes to 8, no-shows fell 40%, and administrative paperwork dropped by three quarters. Every clinical judgment stayed with a clinician, and nothing in the system was allowed to look like one. ## Rule four: every automatic decision is on the record If the system decided something on its own, there is a record of what it decided, on what input, and when. Not because anyone reads it routinely, but because the questions that eventually get asked are retrospective: why did this quote go out at that number, when did we start doing it this way, show me the ten cases before the complaint. An audit trail also happens to be the evidence a regulated market asks for, which is why [the compliance conversation](/en/blog/ai-ethics-responsible-ai) is easier when the operational design came first. Documentation written to describe a system that already behaves well is honest. Documentation written to describe a system nobody constrained is aspiration with a cover page. ## Where the line gets drawn There is no universal answer, and any vendor offering one is selling a poster. | Decision type | Runs unattended | Needs a person | | --- | --- | --- | | Reminders, scheduling, routing, lookups | Yes | No | | Qualification and triage | Yes, with a record | No | | Anything priced, promised, or contractual | No | Always | | Professional judgment in a regulated field | No | Always, by name | | The gray middle | Decided per process | Decided per jurisdiction | The gray middle is where the actual work is, and it gets settled during [the diagnostic](/en/blog/process-audit-before-automation) with the people who carry the liability, before anything is built. A clinic and a brokerage draw the line differently in the same country. The same clinic draws it differently in two countries. Written down before go-live, that line is a design. Inferred afterwards from whatever the system happened to do, it is an incident report. ## What we do not have We do not publish a numbered list of principles, and this post is not one. What exists is the four rules above, with the line redrawn per process during diagnostics. If that sounds less impressive than a framework with six pillars, it is meant to. The useful thing about a rule is that somebody can check on a Tuesday whether it held, and then tell you what happened when it did not. ## FAQ ### Does a human approval step cancel out the speed you promised? It does not, because the slow part of most processes was never the decision. In a brokerage deployment first response went from 6 hours to 8 minutes and document preparation from 2 days to 20 minutes, and a person still signed off on everything that left the building. The six hours were queue time, not thinking time: the inquiry sat in an inbox until somebody got to it. Automation removes the waiting, the retyping, the looking-up and the assembling, and hands a person a finished thing to approve in a minute. That is a completely different use of the approver's attention than asking them to do the work. The cases where approval genuinely slows things down are the ones where the approver is a bottleneck already, and that shows up in the diagnostic as a queue in front of one person. If we find it, we say so, because automating up to a blocked approver just relocates the queue. ### Where exactly should the approval line sit? It is set per process and per jurisdiction, and there is no universal answer worth printing. The rule we apply is that anything binding on the company needs a person: a price quoted, a term agreed, a commitment made, a document that could be read as a contract. Anything purely administrative can run unattended: a reminder, a routing decision, a scheduling slot, a data lookup. The interesting cases sit between, and they get decided during the diagnostic with the people who carry the liability. A clinic and a brokerage in the same country will draw the line differently, and the same clinic in two countries will draw it differently again, because the regulation and the professional duty differ. What matters is that the line is written down before go-live and reviewed when the process changes, rather than being inferred later from whatever the system happened to do. ### What does the human-in-the-loop step actually cost to run? It costs the escalation rate multiplied by the approver's time, and it belongs in the business case as a monthly line item rather than as an assumption. If a process runs 400 times a month, escalates 12% of cases, and each escalation takes an approver 4 minutes, that is roughly 3 hours a month of a specific person's time, every month, forever. That is usually a good trade and it is never zero. Two things make it worse than it needs to be: an escalation that arrives without context, so the approver reconstructs the situation before deciding, and an escalation threshold nobody revisits after launch. We attach the full conversation and the reasoning to every escalation for the first reason, and we review thresholds during handover for the second. When you evaluate any vendor, ask for the escalation rate and the average approver time in the same breath as the payback number. ### How is this different from AI ethics and compliance work? Compliance work answers to a regulator and produces documents: risk classifications, model documentation, bias testing evidence, incident procedures. Human-in-the-loop design answers to the operator and produces behavior: which decision goes where, who is accountable, what gets logged. The two overlap, because auditors ask for evidence that a human review step exists and that it is real rather than nominal, and an audit trail of automated decisions is exactly the evidence they want. But they fail differently. A company can be fully documented and still ship an automation that quietly commits it to a price it did not intend, and a company can have a sound approval design with none of the paperwork a regulated market will demand. Both need doing, and doing the operational half first tends to make the paperwork half honest rather than aspirational. --- # Four Questions That Break Most Automation ROI Numbers URL: https://inite.ai/en/blog/roi-math-for-automation-projects Date: 2026-07-27 Author: Mikhail Savchenko Category: Operations Tags: ROI, Automation, Operations, Procurement ## Direct Answer An automation ROI number is a forecast, and most forecasts break on four questions. Whose hours are being saved, by name and role? Do those saved hours turn into something the company can sell or bank, or do thirty people simply get twenty minutes back each? What volume does the number assume, and what happens at half that volume? And who pays to run and fix the system after go-live? A number that survives all four is worth signing. Payback in our own projects tends to land in three to six months, and it lands there when the saving shows up as capacity that actually gets sold or a cycle that closes faster, rather than as hours counted and multiplied by a salary. ## Key Facts - In one brokerage deployment, lead response fell from 6 hours to 8 minutes and the deal cycle from 14 days to 5. - A private clinic cut patient intake from 45 minutes to 8 and reduced no-shows by 40%. - An equipment rental company moved booking-to-dispatch from 4 hours to 3 minutes and absorbed 2.5x peak-season volume with the same staff. - We put 1-3 workflows into production in 2-4 weeks, and payback typically lands in 3-6 months. - Every build is preceded by a written ROI estimate at conservative inputs; if it does not clear positive, the engagement ends at the diagnostic. ## The number is a forecast Every automation proposal ends with a number. Forty hours a week saved. Payback in four months. A percentage next to the word "efficiency". The number is a forecast produced by the party that benefits from it looking good. That is not a reason to walk away, and it is not usually dishonesty either. Most inflated payback numbers are assembled from pieces that are each defensible and together add up to fiction. Four questions take most of them apart. They work on any vendor, including us, and asking them the same way every time makes answers comparable across proposals. ## One: whose hours, by name? "Saves the team 40 hours a week" is not an answer. Which roles, doing which steps, how many times a week? The reason to insist on names and roles is that the cost of an hour varies by a factor of five inside the same company, and the vague version quietly averages them. An hour of a licensed specialist reviewing something is not an hour of an assistant retyping an address, and a proposal that saves a lot of the second while claiming the rate of the first will look excellent on the page and disappointing in the ledger. Then ask for the loaded cost rather than the salary: employer taxes, benefits, tooling, and the fraction of paid hours that is actually productive. The loaded figure is usually 1.3 to 1.6 times the raw hourly rate, and using the raw one understates the saving. Both errors happen; they just point in different directions and rarely cancel. ## Two: do the saved hours turn into anything? This is the question that removes most of the number, and it is the one almost nobody asks. Twenty minutes a day returned to thirty people is 250 hours a month on a slide and, in most companies, nothing at all in the accounts. Nobody is dismissed, nothing extra is sold, and the time is absorbed by whatever else was already waiting. It is a genuine improvement in how the job feels. It is not payback. Hours become money in two ways, and it is worth naming which one applies before signing: | Route | What has to be true | Example from our deployments | | --- | --- | --- | | Capacity you sell | You were turning work away | Rental firm absorbed 2.5x peak volume without hiring | | A cycle that closes faster | Slow cycles were costing conversions | Brokerage cut the deal cycle from 14 days to 5 | | Headcount you do not add | A hire was actually planned and budgeted | The plan exists in writing before the project starts | The rental example is the cleanest one we have. Booking-to-dispatch went from 4 hours to 3 minutes, and peak-season bookings that used to be refused for lack of turnaround could suddenly be accepted. The hours themselves were incidental. The brokerage case is the second route: response time fell from 6 hours to 8 minutes and the cycle from 14 days to 5, and shorter cycles convert better because fewer buyers cool off in the gap. If neither route applies, the honest description of the project is that it makes the work better, which is a legitimate thing to buy. It should just be bought with that expectation and not against a payback schedule. ## Three: what volume does the number assume? Automation economics are a fixed build cost divided by throughput, so the payback period moves violently with the volume assumption and barely at all with anything else. A process that runs 400 times a month and saves 15 minutes each time is straightforward. The identical process at 80 times a month carries the same build cost against a fifth of the benefit, and it usually fails honest arithmetic even though the per-instance saving has not changed. So: take the volume from your own systems rather than from an interview, use a slow month rather than a good one, and ask for the payback period at half the assumed volume. If the case only survives at the optimistic number, what is on the table is a bet on growth dressed as an efficiency project. That may still be worth taking, but it should be taken knowingly. ## Four: who pays for it after go-live? Build cost is quoted. Running cost usually is not, and it is where four-month payback becomes eleven-month payback. Ask for a monthly figure covering all four of these, then multiply by twelve and [add it to the build cost before dividing anything](/en/answers/ai-automation-cost): - Model and infrastructure per unit of volume, at your real volume. - The people time spent on cases the system escalates, at the honest escalation rate. Any automation that hands uncertain cases to a human is spending that human's time by design, and [keeping a person in the loop for anything binding](/en/blog/process-audit-before-automation) is a deliberate choice with a price attached. - Maintenance when the world moves: a supplier changes a form, a regulator adds a field, a channel changes an interface. - The process owner's continued involvement after handover. An automation nobody owns degrades quietly, and the dashboards keep looking healthy while it happens. ## What a good answer looks like Worked honestly, the arithmetic is short. One process. Its volume from your own system, taken from a slow month. Minutes saved per instance, by role, at loaded cost. A named route by which those minutes become revenue or avoided cost. Twelve months of running cost added to the build cost. Divide. Our own projects put one to three processes into production in two to four weeks, and payback usually lands in three to six months. That range holds because we only sign the ones where the saving has a route to the accounts — which is also why roughly one job in three ends at the diagnostic stage without a build. The [audit that decides which is which](/en/blog/process-audit-before-automation) is deliberately cheaper than the build it might cancel. After go-live, three numbers tell you whether the forecast was real: hours actually freed by named people, the cycle time before and after, and the error rate on the automated step. [Measuring those on a schedule](/en/blog/measuring-ai-roi) is what converts a forecast into a fact, and the schedule should be set before anyone signs. ## Ask us the four questions We hand over the spreadsheet with the inputs visible, because a number the operator cannot interrogate is a number the operator cannot defend to their own board six months later. Bring the four questions to whoever is quoting you next, whether that is us or somebody else. Any vendor who can answer all four in specifics has done the work. Any vendor who cannot has given you a decoration, and [the fastest way to find out which you are holding](/en/blog/business-process-automation) is to ask about volume at half. ## FAQ ### Why should I distrust an ROI number even from a vendor I like? Because the number is a forecast, and the person producing it is the person who benefits from it being attractive. That is not an accusation of dishonesty; it is a description of where the incentive sits. Most inflated payback numbers are not invented, they are assembled out of defensible-looking pieces: an optimistic hourly volume, a headline salary rather than a loaded cost, minutes saved across a wide group that never turn into anything, and a build cost quoted without the running cost that follows it. Each piece is arguable on its own and the total is nonsense. The defense is not skepticism about the vendor, it is a fixed set of questions asked the same way every time, so that the answers become comparable between vendors and between projects. Ask them of us too. If a vendor cannot answer all four in specifics on a call, the number was decoration. ### What is the difference between hours saved and money saved? Hours saved become money in exactly two ways, and it is worth being blunt about which one applies before signing anything. The first is capacity you sell: the same team handles more volume, and the extra volume has revenue attached. An equipment rental company we worked with went from 4 hours to 3 minutes between booking and dispatch and used that to absorb 2.5 times the peak-season volume without hiring, which is real money because peak-season bookings were being turned away before. The second is a cycle that closes faster and therefore closes more often: a brokerage went from a 14-day deal cycle to 5 days, and shorter cycles convert at a higher rate because fewer buyers cool off in the gap. Everything else, particularly twenty minutes returned to thirty people who then do something unmeasured with it, is a real quality-of-life improvement and a fictional line in a business case. Count it as morale, not as payback. ### What running costs do proposals usually leave out? Four, in our experience, and they are the difference between a project that pays back in four months and one that pays back in eleven. First, model and infrastructure cost: every automated decision has a per-unit price, and at real volume that price is a monthly line item, not a rounding error. Second, the human in the loop: any automation that escalates uncertain cases to a person is spending that person's time by design, and the escalation rate has to be estimated honestly rather than assumed to be near zero. Third, maintenance when the surrounding world moves: a supplier changes a form, a regulator changes a field, a channel changes an API, and someone has to notice and fix it. Fourth, the cost of the process owner staying involved after handover, because an automation nobody owns degrades quietly and the reporting keeps looking fine while it does. Ask for these as a monthly figure and add twelve of them to the build cost before dividing. ### How much volume sensitivity should I demand before signing? Ask for the payback at half the assumed volume, and treat the answer as the real answer. Automation economics are dominated by fixed build cost divided across throughput, so the payback period is extremely sensitive to the volume assumption and almost nothing else. A process that runs 400 times a month and saves 15 minutes each time is a straightforward case. The same process at 80 times a month has the same build cost spread across a fifth of the benefit, and it usually fails honest arithmetic even though the per-instance saving is identical. This is the single most common reason a project that looked good on paper disappoints, and it is trivially avoidable: take the volume from your own systems rather than from an interview, use the low month rather than the good month, and ask what happens if the volume never grows. If the case only works at the optimistic volume, it is a bet on growth wearing the costume of an efficiency project. --- # We Built the Same Product Twice. Only 6% of It Carried Over. URL: https://inite.ai/en/blog/inite-estate-real-estate-vertical Date: 2026-07-20 Author: Mikhail Savchenko Category: Operations Tags: Strategy, Automation, Operations, Product ## Direct Answer When a company builds a second product in the same industry, it usually expects to reuse most of the first one. We built a rental platform and then a property platform, which sound almost identical, and only 7 of 113 business concepts turned out to be the same - about 6%. The second product still shipped much faster, because the savings never come from the business logic. They come from the plumbing every product needs and no customer ever sees: accounts, permissions, billing, messaging, notifications, audit trails, translations. Build that once and your second product starts at the interesting part. Expect the business rules to transfer and you will build a compromise that serves neither. ## Key Facts - Of 113 business concepts across our two products, only 7 turned out to be shared - roughly 6%. - The rental product needed 57 industry-specific concepts; the property product needed 63. - Roughly 33% of the second product was infrastructure that had nothing to do with real estate. - The shared foundation covers 8 standard capabilities that every product needs, from billing to audit logging. - We deploy 1-3 automated workflows into production in 2-4 weeks. ## The number that surprised us We build software for two businesses that sound like the same business. One rents things out by the day. [The other sells and manages property](/en/industries/real-estate). Described in a sentence, both are someone paying to use a building or a vehicle for some period of time. When we started the second one, everyone involved assumed most of the first would carry over. Between them, the two products describe 113 business concepts — the things the software has to know about, like a customer, a contract, a price rule, a booking. Seven are shared. Six percent. The second product still shipped far faster than the first. Understanding why is worth more than the number itself, because the same logic decides whether an automation project inside your own company pays for itself. ## Why two similar businesses share almost nothing The sentence that makes them sound alike is the sentence hiding all the differences. A rental company has vehicles. They exist or they do not. A property developer has buildings under construction, where each apartment moves through stages — planned, framed, finished, ready to hand over — and half the business is tracking which stage each one is in. There is no version of a car that is sixty percent delivered, so there was nothing in the first product to borrow. A rental booking opens and closes inside a week. A property sale runs for months and involves a buyer, a seller, an agent, and often a bank, each of whom needs their own view of the same transaction. We know exactly how far you get by treating that as a booking with extra fields: right up until the first commission has to be split three ways. And a rental company has customers. A property company has customers, owners and investors — people who never buy anything through the system and log in only to see what their asset is doing. There is no equivalent at all in the first product, which is the clearest sign that these were never the same business. ## Where the savings actually were Nothing above is where a second product gets cheap. The savings sit underneath, in the part no customer ever sees and no proposal ever itemises. Before a property platform can do anything about property, it needs all of this: | What every product needs | Who notices it | | --- | --- | | Logins, companies, staff roles, permissions | Nobody, until it is wrong | | Billing connected to a real payment provider | Finance, monthly | | Messages from WhatsApp and Telegram in one queue | Whoever answers them | | Notifications, translations, audit trail | Auditors and lawyers | | Encryption of personal data | Everyone, once | That list took months the first time. It is identical whether the message is about a hatchback or a two-bedroom flat. And roughly a third of our second product turned out to be exactly this — infrastructure with nothing to do with real estate. It is the same machinery [any company ends up needing when it automates a process](/en/blog/business-process-automation), which is why it is worth owning once. Build it once and the second product begins where the interesting work begins. That is the whole trick, and it is unglamorous enough that most plans leave it out. ## The rule we use One question decides where anything goes: would a second product do this differently? If two products would do it the same way, build it once. Letting each team write its own version is how a company ends up with four slightly different login systems and fixes the same bug four times. If two products would genuinely differ, keep them separate. A single flexible version covering both cases usually ends up harder to follow than the duplication it was meant to remove. Logins and permissions pass without argument. Pricing fails immediately — seasons and tiers on one side, commissions and multiple parties on the other. When we cannot tell, we keep things separate, because merging later costs an afternoon and separating later costs a quarter. ## Why this matters if you are not building software Most companies that ask us to automate something describe the process as unique to them. The business rules usually are. What is almost never unique is everything around them: pulling inquiries in from several channels into one queue, deciding who handles what, escalating to a person when the system is not confident, recording what happened, and reporting on it afterwards. That machinery is the same for a clinic, a logistics firm and an equipment rental company. It is also the part that takes longest to build from nothing. That is the real reason we can put one to three workflows into production in two to four weeks rather than two to three quarters. Nobody is rebuilding the machinery. You pay for the rules that are yours, which is also why [the payback arithmetic works out at all](/en/blog/measuring-ai-roi). The mistake worth avoiding is the one we nearly made ourselves: assuming that because two things sound alike, the expensive parts will transfer. They rarely do. [Start by finding out which parts of your process are actually yours](/en/blog/process-audit-before-automation), and be honest about how ordinary the rest is — that ordinariness is exactly what makes it cheap. ## FAQ ### If almost nothing transferred, was the shared foundation worth building? It was the only reason the second product was fast. The confusing part is that the savings show up somewhere nobody puts on a plan. Before a property platform can do anything about property, it needs logins, companies, staff roles and permissions, a billing system connected to a real payment provider, a way to receive messages from customers on WhatsApp and Telegram, notifications, translations, an audit trail of who changed what, and encryption of personal data. That list is months of work, it is identical whether you rent out cars or sell apartments, and getting any of it subtly wrong is the kind of mistake that surfaces during a security review rather than a demo. Building it once means the second product starts at the point where the actual business begins. The industry rules were always going to be different - that was never where the money was. ### Why did two such similar businesses share so little? Because they only sound similar. Both are described as someone paying to use a property for a period, and that sentence hides every difference that matters. A rental has a vehicle that either exists or does not; a developer has a building under construction where each apartment moves through stages before anyone can live in it. A rental booking opens and closes in a few days; a property sale runs for months with a buyer, a seller, an agent and sometimes a bank, each needing their own view of the same deal. A rental company has customers; a property company also has owners and investors who never buy anything through the system and log in purely to watch what their asset is doing. Once you write those down, the surprise is not that the overlap was 6%. The surprise is that anyone expected more. ### How do we decide what to build once and what to build per product? The question we ask is whether a second product would do it differently. If two products would do it the same way, build it once, because letting each team write their own version is how a company ends up maintaining four slightly different login systems and fixing the same bug four times. If two products would genuinely differ, keep it separate, because a single flexible version covering both usually becomes harder to understand than the duplication it replaced. Logins and permissions pass that test without argument - a user, a company, who is allowed to do what - and any difference between products there is a bug rather than a feature. Pricing fails it: rental pricing has seasons and tiers, property pricing has commissions and multiple parties, and a universal pricing engine covering both would frustrate everyone. When we cannot tell, we keep it separate, because merging later is easy and separating later is not. ### What does this mean for a company automating its own operations? The same principle applies at a much smaller scale, and it is usually the difference between an automation that pays back and one that quietly does not. Most companies asking us to automate a process describe it as unique to them, and the specific business rules usually are. What is almost never unique is the surrounding machinery: getting messages in from several channels into one queue, deciding who handles what, escalating to a person when the system is unsure, keeping a record of what happened, and reporting on it. That part is the same for a clinic, a logistics firm and an equipment rental business, and it is also the part that takes the longest to build from scratch. When we quote 2-4 weeks for one to three workflows in production, the reason it can be weeks rather than quarters is that we are not rebuilding the machinery each time. You are paying for the rules that are yours, not for the plumbing that is nobody's. --- # Shipping an MCP Server for Your SaaS: A 90-Day Plan URL: https://inite.ai/en/blog/mcp-server-90-day-build Date: 2026-07-07 Author: Mikhail Savchenko Category: Agentic Engineering Tags: AI Agents, Agentic SaaS, MCP, Automation, Strategy ## Direct Answer Customers increasingly want to act through an AI assistant - 'ask Claude to pull my overdue invoices'. To allow that safely, your SaaS needs one standard doorway agents use, built in three 30-day steps. First 30 days: a read-only doorway so an agent can look things up but change nothing, and prove one customer's agent can never see another's data. Next 30 days: allow actions like refunds, but route each through a human who approves first. Final 30 days: harden it with limits, logging and an interface, then test with a live assistant. It is worth doing because the standard (MCP) went from 100,000 to 97 million monthly connections in 18 months - one doorway reaches every assistant at once. ## Key Facts - The standard for connecting AI assistants to software (MCP) grew from roughly 100,000 to 97,000,000 monthly connections in about 18 months. - By early 2026 there were 17,468 public agent-ready servers, up from a few hundred at launch - a market that formed in under two years. - The reference project for this standard passed 84,000 stars from developers by April 2026, one of the fastest-growing of the cycle. - The plan splits into three equal steps of 30 days each: read-only, approved actions, then hardening - 90 days total. - Write actions never run on their own: 100% of them pass through a human approval step before execution. ## The shift worth noticing More of your customers now [start tasks by talking to an assistant](/en/analyze). They ask Claude to summarize a thread, ChatGPT to pull a number, an assistant to book the slot. The natural next step is that they want to do those things *in your product* the same way — and they will lean toward products that let them. For years, supporting that meant a separate custom integration for every assistant, so most companies built none. Then a single standard appeared for connecting assistants to software, and it spread fast: from roughly **100,000 to 97 million monthly connections in eighteen months**, with **17,468 public agent-ready servers** by early 2026. Build one standard doorway now and every assistant — today's and next year's — can use your product. That is the opportunity. Here is a plain 90-day plan to capture it safely. ## Days 1–30: let it look, not touch The first month builds the smallest safe thing: a doorway an assistant can use to *look up* information but not change anything. Pick the one lookup your operators ask for most — "show this month's sign-ups", "list open tickets" — and expose exactly that. Then connect a real assistant like Claude or ChatGPT and confirm two things: it can find and use the lookup, and a request made for one customer cannot see another customer's data. That last point is the entire job of month one. The safe way to guarantee it is simple in principle: the agent's access is tied to one specific customer's credential, so even if the model is talked into asking for someone else's account, the system just refuses. Get that right and the risky months become straightforward. ## Days 31–60: let it act, with a human in the loop Looking things up is safe. Taking actions — a refund, a cancellation — is where the value is *and* where the danger is, so those never run on the agent's word alone. | What the agent can do | When it happens | The safety rule | | --- | --- | --- | | Look things up | Immediately | Scoped to one customer | | Take an action (refund, cancel) | Only after a person approves | Human in the loop | | Bulk or destructive actions | Never through the agent | Staff-only | An action tool doesn't execute — it *proposes*. The proposal lands in a queue, a person on your team approves or rejects it with full context, and only then does it run. The agent proposes; a human decides. It is slower than letting the agent act directly, and it is the only responsible way to give an assistant real power in a live business. It is the same [human-in-the-loop discipline we hold every deployed workflow to](/en/blog/one-engine-many-skins-inite-thesis). ## Days 61–90: make it solid The last month is the difference between a demo and something you run in production: sensible limits per customer, a clear log of who did what, friendly error messages the assistant can understand, and a proper interface rather than raw text. The one piece of advice that saves the most time: don't build a separate agent version of your product. The things an assistant can do should be the exact same capabilities your own in-app assistant already offers, exposed through the one doorway. That way the two can never drift apart, and every capability you already ship [becomes usable by an outside agent for free](/en/blog/mcp-skills-make-saas-ai-native). ## What you have at day 90 A product any major assistant can use, that keeps each customer's data walled off, that never takes a real action without a human's yes, and that logs everything. You didn't bet on which assistant wins — you made your product usable by all of them through one doorway. That is the same reason 17,468 other companies have already done it: build once, and you're present wherever your customers already are. **Related:** [Making your SaaS browser-agent-ready](/en/blog/browser-agent-ready-saas). ## FAQ ### Why should I care about AI agents using my product at all? Because your customers are starting to expect it. People increasingly live inside an assistant - they ask Claude or ChatGPT to summarize, fetch, and book things - and they will prefer products they can drive that way. Historically, supporting each assistant meant a separate, custom integration, so most companies supported none. The change is that there is now one standard doorway (MCP) that every major assistant can use, which is why it grew from about 100,000 to 97 million monthly connections in eighteen months. Building that one doorway means every assistant - the ones that exist now and the ones that ship next year - can use your product with no extra work from you. It is the difference between betting on which assistant wins and simply being usable by all of them. ### Isn't letting an AI touch my system dangerous? It is, if you do it naively - which is exactly why the plan front-loads two safety rules before any risky capability. The first rule is data separation: an agent acting for Customer A must never be able to see Customer B's data, and you guarantee that by tying the agent's access to a specific customer's credential rather than trusting the agent to stay in its lane. A model can be talked into asking for the wrong account; the system simply refuses because the credential only unlocks one. The second rule is that anything which changes or deletes data - a refund, a cancellation - does not run on the agent's say-so. It gets queued as a proposal, and a human on your team approves or rejects it with full context before it happens. With those two rules in place, the worst an agent can do is suggest something a person then declines. That is a very different risk profile from letting it act freely. ### What do the first 30 days actually deliver? A working, safe, read-only doorway - the smallest useful slice. You pick the single most common thing your operators ask for out loud ('show this month's diagnostics', 'list open incidents'), and you expose exactly that one lookup to an assistant. Nothing can be changed yet; it can only read. Then you connect a real assistant like Claude or ChatGPT, confirm it can find and use that lookup, and - most importantly - confirm that a request made with one customer's credential cannot see another customer's data. That last check is the whole point of the first month. It is deliberately small, but it proves the entire approach works end-to-end with a real agent, which is what de-risks everything you add in the following two months. ### How is this different from just having a chatbot? A chatbot talks; an agent-ready doorway lets an assistant actually do things in your product. A typical support chatbot answers questions from a script or a help-center - it cannot pull this specific customer's real invoice or, with approval, issue their refund. Agent-readiness is about giving an outside assistant safe, real access to your live system: look up real data scoped to the right customer, and take real actions that a human has approved. The other difference is reach. A chatbot lives on your site; an agent-ready doorway means the customer can use your product from inside whatever assistant they already prefer - Claude, ChatGPT, or the next one - without you building a bot for each. One is a feature on your page; the other is being present wherever your customers already are. --- # Process Audit Before Automation: The Playbook That Decides When Not to Build URL: https://inite.ai/en/blog/process-audit-before-automation Date: 2026-06-22 Author: Mikhail Savchenko Category: Methodology Tags: process audit, AI automation, ROI, B2B SME, diagnostics ## Direct Answer An audit before an automation build is not a sales asset and not a discovery call. It is a paid five-day diagnostic that produces three artifacts - a swimlane map of the candidate workflow with quantified throughput, error rate and cycle time at every handoff; a cost-of-chaos number in dollars per week; and a written ROI estimate with conservative inputs. If the ROI is not positive at 25th-percentile time-saved and 75th-percentile build-cost assumptions, we refund the diagnostic and the engagement ends. Four patterns explain almost every rejection: wait-dominated bottleneck, mid-rewrite process, adoption blocker, volume too low. ## Key Facts - An engagement can end at the Break-stage diagnostic with no build at all. If the ROI math at conservative inputs does not clear positive, the diagnostic deposit is refunded and the engagement ends there. - The Cut stage removes steps before automation begins rather than encoding them, and every removal is written down with the per-step volume attached — automating a chaotic process is the most expensive way to make slow systems slow forever. - The audit takes five working days end-to-end and produces three named artifacts: a quantified process map, a cost-of-chaos report, and a priority matrix with a written ROI per candidate workflow. Median fee for the diagnostic is 5-10% of the prospective build budget. - Production workflows ship in 2-4 weeks from kickoff, measured against a baseline the operator signed before the build started — which is the only way the after-number means anything. - 4 rejection patterns account for the great majority of audit-stage No's: wait-time-dominated processes where the wait is upstream and not addressable by automation; processes the team itself is mid-rewriting; processes the operators will not adopt due to control or trust constraints; and low-volume processes where the build cost exceeds time-saved over 12 months. ## The rule that defines the company A workflow gets built only after [a written ROI estimate signed by the operator](/en/protocol) clears positive at conservative assumptions. If it does not clear, we refund the diagnostic and the engagement ends. Engagements do end exactly there, and that is the point of having the gate at all — a filter that never rejects anything is not a filter, it is a formality. This is not a sales position. It is the cheapest insurance policy a project can buy. Almost every AI automation engagement that fails six months in failed at the audit stage and the audit did not catch it. ## What the audit is and what it is not The audit is the Break stage of the [INITE Protocol](/en/blog/inite-protocol-6-stages-applied). Five working days. One workflow at a time, sometimes two. Three named artifacts at the end of it. A signed ROI estimate before any build PO is cut. It is not a sales call. It is not a discovery session. It is not a slide deck. The operator pays a small fee for it (typically 5-10% of the prospective build budget) and that fee is refundable in full if the audit ends in a no-build decision. The free fifteen-minute readiness check on the site is a different and earlier thing: it says whether a process is worth auditing at all, and it needs nothing from you but a conversation. This is the step after that one. Once real access to systems and numbers is on the table, the price is the single largest lever on quality — a paid audit gets sent the process owner, a free one gets sent the salesperson. ## The three artifacts | Artifact | What it answers | Time to produce | | --- | --- | --- | | Quantified process map | Where is the actual bottleneck? | 2 days | | Cost-of-chaos report | What does the current way cost in $/week? | 1 day | | Priority matrix + signed ROI | Which candidate workflow has math worth building? | 2 days | The process map is a swimlane picture at handoff resolution — every step from kick-off to closure, each step in its own lane by role, with three numbers stamped on every handoff: throughput per week, error rate, median cycle time. Most teams have never seen their own work in this form. The map is what surfaces the bottleneck, and most of the time the bottleneck is not where the operators thought it was. The cost-of-chaos report is a dollar number per week — what the current way of doing this workflow costs in lost hours, beyond the necessary work. Three lines: rework cost, wait cost, escape cost. Explicit inputs. The company controller signs off on the inputs. The priority matrix is a 2x2 of automation feasibility (technical) and ROI (business) for every candidate workflow surfaced during the audit. The top-right of the matrix is what gets built in the Cast stage. Everything else gets a written explanation of why it did not make the cut. ## The four rejection patterns The audit ends in a No when the math does not clear positive at conservative assumptions. Four patterns produce almost every No. ### Wait-dominated upstream The slow step in the workflow is waiting for something outside the company's control — a regulator's response, a customer's signature, a payment to clear in a bank system. The internal automation would shave minutes off processing, but the cycle time number does not move because the wait is upstream and not addressable. Automating in this case ships a faster horse on a road blocked by a fence. The audit catches this by measuring cycle time at every handoff and labeling whether the wait is internal (operator capacity), inter-team (handoff between functions), or external (vendor, customer, regulator). A workflow with 80% of its cycle time in external waits is a No on a build that proposes to automate internal steps. ### Mid-rewrite The team is in the middle of changing the process for unrelated reasons — an ERP migration, a reorg, a new compliance regime. Automating the current version locks in a process that will not exist in three months. The audit catches this with two questions in the operator interview: what has changed about how you do this in the last 6 months, and what do you expect to change in the next 6? A workflow where the second answer is "everything" or "we are switching systems" is a No until the rewrite settles. ### Adoption blocker The workflow has a trust or control dimension the operators will not delegate. A controller approving an outgoing wire transfer is not going to hand the signing authority to a model, no matter how good the model is. A doctor signing off on a referral is not going to delegate the signature. The automation would technically work and would not be used. The audit catches this by asking the actual operators of the workflow, not their manager, whether they would use the proposed automation. The answer is usually direct. A workflow where the operator says "I would still review every output" is fine if the review takes seconds; it is a No if the review takes as long as the original work. ### Volume too low The candidate workflow happens 8 times a week. Time saved is 90 minutes per instance. That is 12 hours saved per week. Even at $120 loaded cost per hour, the recovery line is $74K per year against a $74K all-in build cost. The recovery line is at or below break-even at year one, and the workflow has to keep running unchanged for two more years to pay back at a respectable rate. That is a No on a build, even though the technology would work and the team would adopt. This pattern is the most common reason a small business asks for an automation that the math will not justify. The audit catches it with the explicit volume × time-saved × cost calculation, and the operator usually agrees with the conclusion the moment they see the numbers laid out. ## The ROI template The math template has three columns and one rule. | Input | How it is taken | Why | | --- | --- | --- | | Time saved per instance | 25th percentile of the operator's observed range | We do not assume the best week | | Build cost over 12 months | 75th percentile of the engineering estimate | We do not assume the smoothest delivery | | Loaded labor cost per hour | Salary + benefits + utilization-corrected | Not raw hourly — the real per-hour cost | The rule — the ROI calculation has to clear positive under those inputs over 12 months. Not 24, not 36, not "eventually." Twelve months. The operator signs the template before any build PO is cut. That signature is what makes the decision shared rather than imposed. A worked example of the arithmetic, on a candidate we did not build. Workflow: inbound RFP triage and routing. Volume: 38 RFPs per week. Time saved per RFP at the 25th percentile: 22 minutes. Loaded labor cost: $95 an hour for a senior account manager. That is 38 RFPs at 22 minutes each, priced at $95 an hour across a 50-week working year, so a recovery line of $66,200 a year. Build cost at 75th percentile: $42,000 all-in. The math cleared by month seven and the workflow shipped. Twelve months in, the realized number was $87K because the time-saved was closer to the median than the 25th percentile, which is the point — conservative inputs, real upside. ## The "40 hours of waste before automating saves 40 hours/week" math The audit's first job is to find out whether the workflow being proposed is the workflow that should be automated. Forty percent of the time it is not. The proposed workflow has a real cost, but the larger cost is in a related workflow upstream, or in the wait between two workflows, or in the rework that comes from a missing input two steps back. This is why the Cut stage that follows the audit removes steps before any automation begins. Steps that should not exist do not get encoded. The audit is what makes that cut visible. A worked pattern. A [logistics-firm engagement](/en/industries/logistics) proposed automating dispatch decision-making. The audit found that dispatchers spent 60% of their time chasing missing documentation from drivers, and only 25% on the decisions the proposed automation would replace. The build pivoted to a document-collection workflow on the driver side. The dispatch decision automation got deferred to phase two and eventually was not needed — the dispatchers had enough capacity once the document chase was gone. That pivot was the audit doing its job. Without the audit, the build would have shipped, the dispatchers would have used it lightly, and the bottleneck would have stayed where it was. ## What "if we cannot show ROI we do not build" actually buys Three things, all of them downstream. Six months after the workflow ships, the operator who approved the build will not be the same person reviewing the results. The original CEO has moved on, or the COO who signed off is at another company, or the operations head is new. The ROI math has to survive that succession. Signed conservative inputs survive. Marketing-flavoured estimates do not. The team that operates the workflow has to trust the system. Operators learn very quickly whether a vendor will refuse work that does not pay back. The vendors who do get a second engagement. The vendors who do not get one engagement and then a polite goodbye. The company's automation capability has to compound. Every workflow that ships against honest math is a workflow whose ROI will be defensible in a budget review. Every workflow that ships against optimistic math gets quietly turned off, and the next automation project gets harder to fund. The audit is the cheapest leverage on whether the next ten projects get green-lit. The CFO-facing version of the same numbers lives in [measuring AI ROI](/en/blog/measuring-ai-roi). ## What this means for an operator considering a build Three operational consequences. First — expect the audit to be paid, expect it to take five days, expect to produce real numbers about throughput and error rate. The price tag is small. The behavior change it produces on the audit team is the largest single lever on quality. Second — expect a 1-in-3 chance that the audit ends in a No. If a vendor's diagnostic-to-build conversion rate is 95%, the diagnostic is sales, not diagnostic. You want the answer to be honest, not flattering. Third - when the audit ends in a Yes, expect the build to ship in 2-4 weeks against signed ROI numbers, with the productivity gain measured against a baseline that you signed before the build started. A gain measured that way is real on the workflows that survive the filter. It is not real, in any sense worth paying for, on the workflows that should never have been built. The audit is what keeps the difference visible. The single-process discipline that follows the audit is in [AI integration in business](/en/blog/ai-integration-business). The four questions an operator should ask about any payback number, ours included, are in [four questions that break most automation ROI numbers](/en/blog/roi-math-for-automation-projects). ## FAQ ### If you refund the diagnostic when the math says no, what stops you from just rubber-stamping the math? Three guards. First, the math itself has conservative defaults written into the template — time-saved is taken at the 25th percentile of the operator's own week-on-week range, build cost is taken at the 75th percentile of the engineering estimate, and labor cost is loaded (salary plus benefits plus utilization-corrected, not raw hourly). A workflow has to clear positive ROI under those bad assumptions, not under the optimistic ones. Second, the ROI calculation is a signed artifact — the operator sees the numbers, the inputs, and the assumptions before any build PO is cut. They are part of the rejection decision rather than a recipient of it. Third, the rejected engagements get a written one-page memo explaining which of the four patterns they fell into; that memo is filed and tracked. The memo is what makes a rejection auditable afterwards, and it is the only evidence that the filter does real work rather than theatre: a vendor who cannot produce one has never rejected anything. ### What does the swimlane map produced at the audit actually look like? A picture of one workflow at handoff resolution — every step from kick-off to closure, each step in its own lane by role, with three numbers stamped on every handoff: throughput per week, error rate (rework or correction), and cycle time (median wait between handoffs). Most teams have never seen their own work in this form. The map is what surfaces the bottleneck, and most of the time the bottleneck is not where the operators thought it was. We use BPMN notation when the team is already used to it; plain rectangles and arrows otherwise — the notation is not the point. The point is that the three numbers are visible per handoff, so the conversation about where to automate is grounded in data rather than the loudest voice in the room. The map plus the three numbers per handoff are typically two business days of work, including the interviews and the observed shadowing. ### What is the cost-of-chaos report and how is it actually computed? A dollar number per week — what the current way of running this workflow costs the company in lost hours, beyond the necessary work. It is the sum of three lines. Line one — rework cost — error rate per handoff multiplied by the median rework time multiplied by the throughput multiplied by the loaded hourly cost of the role doing the rework. Line two — wait cost — median cycle time per handoff multiplied by throughput multiplied by the loaded cost of the resource being held up (a sales rep waiting on credit approval costs more than an accountant waiting on a signature). Line three — escape cost — the dollar value of work that left the funnel because of the bottleneck (lost deals, refunded jobs, customer escalations resolved by credits). All three lines have explicit inputs, all three are visible in the artifact. The number is conservative and the company controller signs off on the inputs. This is the number against which build ROI gets measured. ### Which audit findings most often lead to a No? Four patterns account for most rejections. Wait-dominated upstream — the workflow's slow step is waiting for an external party (a regulator, a customer signature, a payment to clear), and automating any internal step does not move the cycle time because the wait is outside the company's control. Mid-rewrite — the team is in the middle of changing the process for unrelated reasons (an ERP migration, a reorg, a new compliance regime); automating the current version locks in a process that will not exist in three months. Adoption blocker — the workflow has a trust or control dimension that operators will not delegate (a controller approving an outgoing wire, a doctor signing off on a referral); the automation would technically work and would not be used. Volume too low — the candidate workflow happens 8 times a week, time saved is 90 minutes per instance, that is 12 hours saved per week against a 6-month build; even at $120 loaded cost the recovery line never crosses. Each of these gets a written memo, the operator sees which pattern they hit, and the engagement either pivots to a different candidate workflow or closes. ### Why a paid diagnostic? Why not just audit for free as part of sales? First, a definition, because 'diagnostic' covers two different things here. The fifteen-minute readiness check is free and stays free — it needs nothing but a conversation and it only answers whether a process is worth auditing. This question is about the five-day audit, which needs access to your systems and your real numbers. Two reasons that show up in operator behavior the moment money is on the table. Free diagnostics get treated as sales meetings — the operator sends the BD-friendly person, the access to real numbers stays gated, and the audit becomes an exchange of vendor-comparison answers rather than a process investigation. Paid diagnostics get treated as work — the operator sends the actual process owner, opens the real systems, and answers the boring questions about throughput and error rate honestly. The price tag of the diagnostic is small (typically 5-10% of the prospective build budget) but the behavior change is the largest single lever on diagnostic quality. The refund clause makes the price psychologically reversible: if the math does not work, the operator's downside is the time they spent, not the cash. That trade — small fee for honest access, refundable on no-build — produces audits that actually inform the decision rather than diagnostics that justify the decision the salesperson already wanted. --- # INITE Protocol: 6 Stages Applied to a Real Q2-2026 Deployment URL: https://inite.ai/en/blog/inite-protocol-6-stages-applied Date: 2026-05-25 Author: Mikhail Savchenko Category: Methodology Tags: INITE Protocol, AI Automation, Process Audit, B2B SME, ROI ## Direct Answer The INITE Protocol is a 6-stage transformation methodology - Break (Week 1-2, diagnose and reset), Hold (Week 2-3, stabilize critical processes), Track (Week 3-4, measure and analyze), Cut (Month 2, simplify and eliminate), Cast (Month 2-3, build and deploy 1-3 production workflows), Form (Month 3-6, optimize and scale). The first automated workflow goes live in 2-4 weeks; the full cycle runs 3-6 months. The bottleneck has never been the technology - it is whether the diagnostic finds a process the math says is worth automating. If the math does not clear at conservative inputs, we do not build, and the engagement ends at the diagnostic. ## Key Facts - 6 sequential stages over a 3-6 month full cycle: Break (Week 1-2), Hold (Week 2-3), Track (Week 3-4), Cut (Month 2), Cast (Month 2-3), Form (Month 3-6). First production workflow live in 2-4 weeks. - Each stage has to hand the next one a named artifact: Break a cost-of-chaos model, Hold a stabilised process map, Track a signed baseline, Cut an elimination log, Cast a running workflow, Form a monitoring dashboard. - The Break stage is allowed to end the engagement: if the ROI math at conservative inputs does not clear positive, the diagnostic deposit goes back and nothing is built. - The Cut stage removes steps rather than encoding them, and every removal is written down with the per-step volume from Track attached as evidence - automating the chaos is the most expensive way to make slow systems slow forever. - The Cast stage ships 1-3 production workflows, not pilots. On this deployment the first workflow went from spec-frozen to live with real leads in 8 calendar days. ## What this post is, and what it is not The INITE Protocol is the 6-stage methodology behind every B2B SME deployment INITE ships. Break, Hold, Track, Cut, Cast, Form. The names are stage labels and most readers can guess the meaning from the word. What is harder to convey - and what every consultancy-flavoured methodology overview leaves out - is what the team actually does at each stage, what artifacts come out, and which decisions are made. This post fills that gap. It walks one Q2-2026 deployment at a 60-person professional-services firm, twelve of them fee-earning consultants (anonymized; sector, scale and process pattern preserved, names and commercially sensitive figures removed). Two production workflows shipped in 19 calendar days from the Cast kick-off, against a $24K build. Year one roughly washes. Year two is where the return sits. That is the shape a first automation actually has, and it is the shape most case studies decline to show. Every number below belongs to this one engagement, and every one of them derives from another number on this page: volumes come out of the systems, hours out of the volumes, money out of the hours at a stated loaded rate. Where a benefit could not be priced that way it is named and left out of the return rather than estimated into it. INITE publishes no cross-client average, because there is no measured one to publish. The protocol itself is documented in `lib/brand-canonical.ts` (`whatShipped`), `locales//common/protocol.json`, and the HowTo JSON-LD in `components/StructuredData.tsx`. The discipline behind the Break stage in particular is in [process audit before automation](/en/blog/process-audit-before-automation). ## Stage 1 — Break (Week 1-2): Diagnose & Reset The Break stage answers one question: is there a process here whose automation math survives scrutiny? If yes, we keep going. If no, we end the engagement and refund the diagnostic deposit. **What we do.** Stakeholder interviews (CEO, COO, 1-2 frontline operators - in this case the partner running the practice, the office manager, and 2 senior consultants). Process observation - we shadow the actual work for 4-8 hours per target workflow. Data pull - we ingest 90 days of operational data from the existing systems (CRM, project tool, billing, time tracking). Baseline measurement - we instrument the current process with timing, throughput, and error counts. **Artifacts produced.** 1. **Process map with bottleneck markers.** A swimlane diagram of every step in the target workflow, with quantified throughput, error rate, and cycle time at each handoff. For this deployment we mapped 3 candidate workflows. **(a) Inbound lead qualification:** 34 leads a week, 22 minutes of ops-coordinator time each for research, CRM entry and routing, so 12.5 hours a week, plus 6 hours a week of consultant time redoing work on mis-routed leads. Median first response 38 hours; 24% of leads had no reply inside five working days. **(b) Project status updates:** 26 active engagements, one note a week each, 40 minutes of consultant time per note, so 17.3 hours a week. **(c) Invoice and time-entry reconciliation:** 8 hours a week of office-manager time at a 22% error rate. 2. **Cost-of-chaos report.** The hours above, priced at loaded cost: $95 an hour for a consultant, $45 for the ops coordinator and the office manager, which is what the hour costs the firm to provide rather than what appears on a payslip. Lead qualification $1,133 a week, status updates $1,647, reconciliation $360. Total **$3,140 a week**, annualised over 48 working weeks rather than 52, because holiday and slack weeks do not get automated away: **$151K a year**. That is the number every later claim gets measured against. The 38-hour first response and the 24% of leads that never got a reply are not in it. Both are real, and both are probably worth more than the hours, but converting them into money means assuming a conversion rate and a deal value, and an assumption is not a measurement. They stay on the page as findings and out of the return. 3. **Priority matrix.** Every candidate workflow scored on automation feasibility (technical, 0-100) against return (business, 0-100). The top of the matrix is what gets built in Cast; the bottom is explicitly deferred. | Workflow | Feasibility | ROI | Decision | | --- | --- | --- | --- | | Inbound lead qualification | 82 | 88 | Build in Cast (W1 of Cast) | | Project status updates | 71 | 76 | Build in Cast (W2 of Cast) | | Invoice / time reconciliation | 54 | 38 | Defer - requires data-quality work in Hold first | Lead qualification scores higher on return than its hours alone justify - it is the smaller of the two lines at $1,133 a week against $1,647. The score is the one place the unpriced drop-off is allowed to count, because the matrix is a ranking rather than a sum and a ranking can carry a judgment. The money below cannot, which is why the same drop-off stays out of every figure in the rest of this post. Saying which of the two is happening is the whole difference between a matrix and a guess. The third workflow does not get built in this engagement. It is documented as a follow-up candidate for a future cycle, but only after the Hold stage cleans up the data-quality issues that make it both lower-feasibility and lower-return today. Honest deferral is part of the methodology - we do not bundle the bad one in to keep the engagement bigger. **Decision at the end of Break.** The two workflows worth building carry $2,780 a week of the $3,140, which is $133K a year. The model assumed automation would take back 60% of those hours rather than all of them, because the exceptions stay human: **$70-85K a year**. Build estimated at $24K, monitoring and tuning at $750 a month afterwards. The workflows go live at the end of month 3, so year one collects nine months of benefit: roughly $52K against $33K of cost, a net near **+$19K with payback around month 9**. Year two, with only the support cost left, sits near +$61K. Decision: proceed to Hold. Nobody in that room found this exciting, which is the correct reaction and the reason it was approved. A first automation that pays for itself inside a year and compounds afterwards is a normal good outcome. A first automation projected to return four times its cost in twelve months is a sales document. A Break stage is allowed to end here with a no. Had the arithmetic at conservative inputs not cleared - and the third workflow on its own would not have - the diagnostic deposit goes back and the engagement stops. That has to be a real outcome or none of the measurement above means anything. ## Stage 2 — Hold (Week 2-3): Stabilize & Fix The Hold stage answers: can the target process be automated in its current state, or does it need stabilization first? Automating chaos is the most expensive way to make slow systems slow forever; this stage stops that. **What we do.** For each workflow that will be built in Cast, we identify the upstream data sources, validation rules, and handoff points that must be reliable for automation to work. We fix the ones that are broken. We document the SOPs (standard operating procedures) for the manual versions of the steps that will remain human - because the new automated workflow will still hand off to humans at the edges, and the human side needs to be consistent. **Artifacts produced.** 1. **SOPs for key workflows.** For this deployment: the email triage rules for inbound leads (what counts as a qualified lead, what gets escalated, what gets discarded - written down for the first time); the status-update template the team had been improvising; the data field requirements for the CRM (which fields are mandatory at lead capture vs at conversion). 2. **Data quality fixes.** The CRM had three lead-source values where there should have been one (`"website"`, `"web"`, `"Website"`). We collapsed them. The project tool had 4 active status values where the team only used 3. We removed the dead one. The email inbox had 14 inbox rules accreted over 5 years, half of them dead. We pruned to 6 active rules. 3. **Quick manual fixes that save time immediately.** Two of the bottlenecks in the lead-qualification workflow turned out to be fixable without any AI - one was an email forwarding rule that misrouted leads to a vacation auto-responder; the other was a CRM required-field bug that forced consultants to re-enter the same data twice. Both were fixed in Hold, before Cast started. The Hold-only time savings ran about 6 hours a week across the team - a small win before any automation shipped, and those hours sit inside the week-12 figures below rather than on top of them, because the same hour cannot be counted twice. Hold is the stage that most "AI pilot" methodologies skip, and it is the stage that determines whether the eventual Cast workflow has stable inputs to work against. An LLM-driven workflow that receives inconsistent lead-source values, ambiguous status values, and emails routed to the wrong inbox will produce garbage and the team will blame the AI. The Hold stage ensures the substrate is clean. ## Stage 3 — Track (Week 3-4): Measure & Analyze The Track stage answers: what does the process actually look like in instrumented production, not in the interview-based map we drew in Break? **What we do.** Instrument every step of the workflow with timestamped event logging - typically a lightweight middleware that records process events (lead received, lead qualified, lead converted, status update sent, etc.) with timing, identity, and outcome. Collect 2-3 weeks of data on the now-stabilized process from Hold. Find patterns the static map missed. **Artifacts produced.** 1. **KPI dashboard (live metrics).** A real-time dashboard of the per-workflow KPIs that will be tracked through Cast and Form. For this deployment: per-workflow cycle time, throughput per consultant, intervention rate (how often the operator overrides an automated decision), and error rate. The dashboard goes live in Track and stays live through every subsequent stage - the team sees it daily. 2. **Pattern analysis report.** The non-obvious findings from the instrumented data. For this deployment, three findings stood out. (a) 31% of inbound leads came in between 6pm and 8am local time - the manual workflow had nobody on shift then and these leads sat for 12-16 hours before first response (a fact the team had not measured because nobody was watching during off-hours). (b) Most status updates went out in a single batch on Friday afternoon, which is why clients experienced the reporting as bursty rather than weekly, and why nobody inside the firm had noticed - from the inside it feels like a weekly rhythm. (c) The two consultants who wrote the longest status updates were also the two with the highest renewal rate on their engagements. Two consultants is not a sample and it proves nothing; it was still enough to decide against tuning the automation toward brevity, which is the kind of call a dashboard cannot make for you. 3. **Automation feasibility scores per process step.** A line-by-line evaluation of which steps in each workflow are good automation candidates (high-volume, well-defined inputs, clear success criteria) vs which should stay human (low-volume, ambiguous inputs, judgment calls). For this deployment, lead qualification scored 82/100 (highly automatable), the first-draft status-update generation scored 76/100 (automatable with human-in-loop review), and the final client-facing send scored 38/100 (should stay human, even after AI drafts). Track is what transforms the Break-stage hypothesis ("this process looks slow and error-prone") into Cast-stage specification ("this workflow has these specific bottlenecks at these specific steps, with this specific data shape"). A team that skips Track ships a Cast that fixes the wrong thing. ## Stage 4 — Cut (Month 2): Simplify & Eliminate The Cut stage answers: which steps in the process should not exist at all, before any of them get automated? **What we do.** Take the now-instrumented process map from Track and walk it step by step. For each step, ask: does this step add value, or is it scar tissue from a problem that no longer exists? Eliminate the ones that do not. Merge duplicate steps. Re-order steps where the order is arbitrary. Prepare the cleaned process as the specification for Cast. **Artifacts produced.** 1. **Streamlined process flows.** The cleaned-up version of each target workflow, ready for automation. For this deployment, the lead-qualification workflow went from 11 steps to 6. Five eliminated steps: one duplicate data entry (CRM was being updated twice for legacy compatibility reasons that had not been true for 18 months), one redundant manager approval (added during a previous quality issue that had been resolved), two status-tracking emails that nobody read (we measured open rates - 4% and 7%), and one manual data export that fed a report nobody ran anymore. 2. **Eliminated redundant steps - count and categories.** Across the 2 target workflows, 9 steps eliminated of 22 - 41%. That is a high proportion, and the reason is specific rather than typical: the firm had not done a process review in 4 years and the scar tissue had accumulated. Categories: 4 duplicate data entries, 2 redundant approvals, 2 dead notifications, 1 manual feed to a dead report. 3. **Automation-ready specifications.** For each step that remains and is to be automated, a written spec - input data shape, success criteria, error handling, escalation path, monitoring metrics. This is the artifact Cast builds against, and the test of its length is whether the operations lead will read the whole thing in one sitting rather than skim it. Both specs here passed that test. Written specs prevent the most common Cast failure mode, which is the engineering team building something that does not match what the operations team agreed to in Track. Cut is the stage that most CTO-led "let's add AI" initiatives skip because nobody wants to argue with the team about eliminating their pet step. We argue. The math wins; the elimination decisions are written down with the per-step volume from Track attached as evidence. A team that does not Cut ends up automating the chaos and locking it in. ## Stage 5 — Cast (Month 2-3): Build & Deploy The Cast stage answers: do the spec'd workflows ship to production, and do they survive contact with real users? **What we do.** Build each workflow against the specification frozen in Cut. Ship to production - not a demo environment, not a sandbox, the real environment with the real users. Wire monitoring before launch. Integrate with the existing tools the team already uses - the CRM, the inbox, the project tool - rather than replacing them. For deployments that sit on the Inite ecosystem (see the companion post on the [shared @inite/* runtime](/en/blog/one-engine-many-skins-inite-thesis)), the Cast stage reuses `@inite/assistant` for the LLM runner, `@inite/inbox` for any conversation surface, `@inite/api-kit` for the request-wrapper pattern, `@inite/incidents` for the human-in-loop escalation path, and `@inite/security` for PII masking on the audit log. The reused infrastructure compresses Cast from "build everything" to "configure most of it and write the domain logic". For this deployment, the lead-qualification workflow shipped in 8 calendar days from spec-frozen to live-with-real-leads. **Artifacts produced.** 1. **1-3 automated workflows in production.** For this deployment: (a) lead-qualification workflow live in CRM + email inbox - 100% of inbound leads now flow through the AI triage, with operator-override available at any step; (b) project status-update workflow live in the project tool + email - generates draft status updates which the consultant reviews and sends. Both workflows hit production within 19 calendar days of Cast start. 2. **Integration with the firm's existing tools.** No new dashboard for the team to learn. The lead workflow appears as automation rules and AI-generated comments inside the existing CRM. The status-update workflow appears as drafts in the existing email-compose flow. The team's daily tools are unchanged; the AI is invisible plumbing inside them. Adoption is the silent killer of pilots; building inside the team's existing tool set is how Cast avoids it. We use `@inite/assistant`'s tool registry pattern to expose the workflows to any agent that needs to call them - including the [MCP-compatible agents the firm's CTO uses for code review](/en/blog/mcp-skills-make-saas-ai-native). 3. **Team training and handover.** A 90-minute live session per workflow with the team that will operate it, plus a written runbook (3-5 pages) for each workflow covering: how the workflow normally runs, what monitoring metrics to watch, what to do when an alert fires, who owns the workflow internally, how to escalate to us. The runbook is the bridge to Form. Cast is also where the [browser-agent-readiness checks](/en/blog/browser-agent-ready-saas) get applied to any operator-facing surface the workflow exposes - because if an AI agent is going to drive the firm's tools to assist humans, the tools themselves need to be agent-readable. This is a small but increasingly important detail in 2026 deployments. ## Stage 6 — Form (Month 3-6): Optimize & Scale The Form stage answers: does the deployed workflow survive 12 months of real usage, and does the team accumulate a capability instead of getting a one-off project? **What we do.** Watch the production traffic for the first 90 days. Tune the workflow against real data, not assumed data. Scale to additional workflows on the same platform. Hand over operational ownership to an identified internal owner. **Artifacts produced.** 1. **Performance monitoring dashboard.** The KPI dashboard from Track, now wired to the production workflows and watched daily by the internal owner. Week-12 measurements, against the baseline the operator signed at the end of Track: median first response on an inbound lead from 38 hours to 1.4 hours; leads with no reply inside five working days from 24% to 9%; ops-coordinator time on lead handling from 12.5 to 3.5 hours a week; consultant time on status updates from 17.3 to 5.2 hours a week; consultant rework on mis-routed leads from 6 to 1.5 hours a week. Exceptions still go to a person, which is why none of those figures goes to zero and why a proposal that promises one should be read carefully. 2. **Optimization based on real usage data.** Four rounds of tuning in the first 90 days. Week 3: the classifier was too willing to mark borderline leads unqualified, so the threshold moved and the ops coordinator stopped having to re-check the discard pile every morning. Week 6: the drafts read generically, so per-client context retrieval from the project tool went in. Week 9: the escalation path was noisy enough that people had begun ignoring it, so a confidence gate went in front of it. Week 12: the override log showed three lead types that should never auto-qualify, and those became hard rules. None of this could have been done before go-live, because none of the data existed. 3. **Scaling plan for next automation wave.** With the platform live, the team's bandwidth opens up. We documented 4 follow-on candidates for the next cycle: the invoice/time-reconciliation workflow deferred in Break (now feasible because Hold cleaned the data); a client-onboarding automation; a proposal-drafting assistant; a quarterly-business-review automation. Each scored on the same feasibility × ROI matrix used in Break. The firm picked 2 to run in Q3-2026; we are scoping those as a separate engagement. The handover is the hardest part of Form. The internal owner gets root access to the workflow config, the monitoring dashboard, the prompt registry, and the escalation runbook. We stay available on a quarterly check-in for the first year, but the operational responsibility is theirs. Workflows that do not have an identified internal owner with budget for tuning are the ones that quietly go dark in 6 months. We refuse to ship Cast without Form, and we refuse to call Form complete without the internal owner. ## What the deployment cost, and what it produced | Metric | Baseline (Break) | Week 12 | Change | | --- | --- | --- | --- | | Median first response, inbound lead | 38 hours | 1.4 hours | −96% | | Leads with no reply inside 5 working days | 24% | 9% | −15 pp | | Ops-coordinator time, lead handling | 12.5 hr/week | 3.5 hr/week | −9 hr/week | | Consultant time, status updates | 17.3 hr/week | 5.2 hr/week | −12.1 hr/week | | Consultant rework, mis-routed leads | 6 hr/week | 1.5 hr/week | −4.5 hr/week | | Hours reclaimed, at loaded cost | — | $1,982/week | $95K/yr over 48 weeks | | Build, plus first-year support | — | $33K | $24K build, $750/month | | Year-1 net, benefit from month 4 | — | +$38K | payback in month 8 | | Year-2 net, support cost only | — | +$86K | — | The reclaimed hours came in at $95K a year against a modelled $70-85K. That is the direction an estimate should miss in, and it is the reason the model assumed 60% recovery rather than 90%. The drop-off improvement is not in any of those figures. A quarter of inbound leads going unanswered, down to under a tenth, is at this firm's deal sizes almost certainly worth more than every hour in the table put together. It stays out because nobody measured what those leads would have converted at, and a number built on an assumed conversion rate would have made the whole table unfalsifiable. If the firm instruments that next year, it becomes a measurement and it goes in. The numbers are for this deployment and nothing else. There is no published INITE average to hold them against, and there should not be one until it has been measured across a sample worth quoting. The CFO frame for projecting a lift like this onto the P&L is in [measuring AI ROI](/en/blog/measuring-ai-roi). ## What the protocol is, in one sentence [The INITE Protocol is what a methodology looks like when it is built backwards](/en/protocol) from the question "did the workflow survive 12 months of real usage?". Break / Hold / Track ensure the math is real. Cut ensures the substrate is worth automating. Cast ships production-grade software against a clean process. Form makes the change stick. Skipping any one of the six is how AI consultancy money goes to die. Following all six is how a workflow is still running the quarter after we leave. If your math survives Break, your team owns Form. Everything between is engineering. The substrate that makes the Cast stage compress from 3 weeks to 8 days is described in [the Inite operating model for vertical AI products](/en/blog/one-engine-many-skins-inite-thesis). ## FAQ ### Why 6 stages instead of just 'audit then build'? Because audit-then-build is the most expensive way to fail. Two failure modes show up every time: (1) the audit identifies a real bottleneck but the chosen automation does not actually fix it because the process is non-deterministic upstream - we automate a downstream symptom and the bottleneck moves; (2) the build ships software that works but no one uses it, because the process around it is unchanged and the team has no reason to switch. The 6 stages prevent both. Break / Hold / Track create a measured baseline. Cut eliminates the steps that should not exist before automation locks them in. Cast ships production-grade software against a clean process. Form makes the change stick. Skipping any of the 6 trades short-term speed for the long-term certainty that the workflow will be quietly abandoned in 6 months. ### What does Stage 1 - Break actually produce? Three artifacts. (1) A process map with bottleneck markers - typically a swimlane diagram of every step in the target workflow with quantified throughput, error rate, and cycle time at each handoff. We use BPMN notation when the team already knows it; otherwise plain rectangles with arrows. (2) A cost-of-chaos report - the dollar value of hours lost per week to manual rework, missed handoffs, and waiting. This is the number against which ROI gets measured. (3) A priority matrix - every candidate workflow ranked by automation feasibility (technical) × ROI (business). The top of the matrix is what gets built in Cast; the bottom is what gets explicitly deferred. If no candidate has positive ROI even at conservative time-saved assumptions, we end the engagement and refund the diagnostic, and that outcome has to be available or the other two artifacts mean nothing. ### How is Stage 5 - Cast different from a typical 'AI pilot'? Three differences. (1) The output is 1-3 workflows in production, not a demo environment - same auth, same data, same operators, same SLA as the rest of the company's stack. (2) Each workflow ships with monitoring wired before launch - latency, error rate, intervention rate (how often a human override is needed), and the per-workflow KPI from the Cut stage. We track these from day one, not after the launch dust settles. (3) The build sits on top of the existing tools the team already uses - we wire the AI into the CRM, the inbox, the spreadsheet, the chat tool - we do not replace them. Adoption is the silent killer of pilots; building inside the team's existing tool set is how Cast avoids it. ### What happens in Stage 6 - Form that does not happen in Stage 5 - Cast? Cast ships the workflow live. Form makes it survive 12 months. Three things happen in Form. (1) Tuning against real usage data - the prompts, retrieval thresholds, escalation rules, and routing logic get adjusted based on what the production traffic looks like, not what the spec assumed. Typically 3-5 rounds of tuning in the first 90 days. (2) Scaling to the next 1-2 workflows on the same platform - the second workflow takes about 40% of the time of the first because the infrastructure (auth, monitoring, agent runtime, prompt registry) is reused. (3) Handover with documentation, runbooks, and an internal owner identified - we are not the long-term operator of the workflow; the team is. Form is what makes a deployment a capability instead of a one-off project. Skipping Form is how the workflow gets quietly turned off in 6 months when the original champion leaves. ### How does the protocol interact with the wider Inite vertical AI ecosystem? The protocol is product-agnostic on its face - Break / Hold / Track / Cut / Cast / Form would work for any B2B automation engagement. In practice, when a deployment ships on top of the Inite ecosystem (the shared @inite/* runtime described in the companion thesis post), the time costs compress significantly: the Cast stage reuses @inite/assistant for the LLM runner, @inite/inbox for any conversation surface, @inite/api-kit for the request-wrapper pattern, and @inite/incidents for the human-in-loop escalation path. A workflow that would take 3 weeks to build from scratch typically ships in 8 calendar days when it sits on the shared runtime. The protocol stays the same; the substrate is what makes Cast cheap. ### What does 'if we cannot show ROI we do not build' mean in practice? It is the rule that defines the company. Before any build starts, the Break stage produces a written ROI estimate with three inputs: time-saved per process instance × instances per week × loaded labor cost per hour, minus the all-in build cost over 12 months (engineering + monitoring + tuning). If that number is not positive at conservative assumptions (we use the 25th-percentile estimate for time-saved and the 75th-percentile estimate for build cost), the engagement ends at the diagnostic and we refund the deposit. The point is not to be picky - it is to make sure that every shipped workflow has math that survives scrutiny six months in, when the original CEO who signed off has moved on and the new operations head is asking what this thing costs. --- # One core, eighteen entries in the ledger and one in production URL: https://inite.ai/en/blog/one-engine-many-skins-inite-thesis Date: 2026-05-18 Author: Mikhail Savchenko Category: Architecture Tags: Architecture, Multi-tenant, Strategy ## Direct Answer Our industry products stand on one shared base: five mandatory entities, nineteen capability packages and three services - sign-in, invoices and assistant. Written once, never rewritten per industry. Only the subject matter is written fresh, and its share is larger than people expect: in a 318,757-line rental CRM, 88% of the data schema was industry-specific, and across two of our products in adjacent industries 7 of 113 business concepts were shared. The ledger currently holds 18 entries, 1 of them in production and 2 in pilot. We publish the ledger of statuses rather than a count of products, because only the count can be wrong in a way nobody notices. ## Key Facts - The ledger holds 18 entries, 1 in production and 2 in pilot, and not one has reached full conformance with the specification. - There are 5 mandatory entities: user, company, the role between them, an access key, and a permission override for one company. - There are 19 capability packages and 3 horizontal services: sign-in, invoices and assistant. - In a rental CRM of 318,757 lines, 1650 of 1872 schema lines were industry-specific, that is 88%. - Of 113 business concepts across two of our products in adjacent industries, 7 were shared, about 6%. ## What sits under all the products at once [In any system with staff and customers the same things repeat](/en/industries). Someone signs in under their own account. They hold a role. The role permits some things and forbids others. Notifications go to someone, an invoice to someone else, the correspondence lives in one place and is findable, and a log remembers who changed what. That is five mandatory entities - user, company, the role between them, an access key, and a permission override for an individual company - plus nineteen capability packages and three services: sign-in, invoices and assistant. Written once, never rewritten for an industry. The fifth entity did not come from a plan. It came from a case. A branch manager in the rental business needed access to reporting meant for fleet co-owners, and an ordinary manager role does not grant it. The role table could not express that. The trigger was industry-specific; the fix turned out to be general, and every product now inherits it. ## How much actually transfers The custom here is to quote a large share. We measured twice, and both figures were inconvenient. The rental CRM we moved onto the shared base is 318,757 lines and four years of accumulated rules. Of 1872 schema lines, 1650 were industry-specific, that is 88%. The shared scaffolding was 220 lines. The second measurement is harsher. We built a rental platform, then a real-estate platform. Adjacent industries, both about an object handed to somebody for a while or for good. Of 113 business concepts across the two products, seven were shared. [About 6%](/en/blog/inite-estate-real-estate-vertical), and that number is worth reading about on its own. Which yields a consequence useful well outside our own kitchen. "We already have this, we just need to configure it" is a sentence about scaffolding, not about your business. [Where four weeks comes from](/en/blog/4-week-vertical-cloning-playbook) works through the same thing from the side of the person choosing a vendor. ## What backs this up The rental move took four weeks, and operators were working in the rebuilt application from the third, running real bookings through it. The bottleneck was not build time. Three times it turned out that the real rental data contradicted the specification of the base, and each time we rewrote the specification rather than the data. A base that cannot survive contact with a four-year-old working product is not a base yet. One side effect mattered more than anything planned. Business rules such as "no vehicle handover until the contract is signed and the deposit is in" went straight into the layer that external AI agents call. The agent now cannot get around the rule without even knowing it exists: the layer returns a refusal naming the condition it broke. Every product after that gets it for free. ## The scoreboard, unflattering | Status | Count | | --- | ---: | | In production | 1 | | Pilot | 2 | | Awaiting migration | 4 | | Drifting from the specification | 4 | | Not brought onto the base | 2 | | Named, not started | 3 | | Replaced by another | 2 | | **Full conformance with the specification** | **0** | Eighteen entries. One in production. And the status the specification names as the goal has not been reached by any of them yet. A product count is a marketing number, a ledger of statuses is an engineering one, and only the first can be wrong in a way nobody notices. "Eighteen products on a shared base" outlives any drift underneath it. "One in production, none in full conformance" creates work. For anyone building a family of products of their own, what transfers from here is not our code. It is the habit: decide before the first product which layer is allowed to have opinions, put the shared base behind a version, and keep the ledger honest enough to be unpleasant to look at. A ledger nobody winces at is a ledger nobody updates. What this looks like from the customer's side rather than the builder's is worked through in [what to automate first](/en/blog/what-to-automate-first-in-a-small-company). ## FAQ ### What exactly is shared, and what gets written again every time? Shared is whatever works the same way in any system with staff and customers: who signed in under their own account, what role they hold, what the role permits and forbids, who gets notifications, who gets invoiced, where the correspondence lives and who changed what. Encryption of personal data and translations belong here too. None of it depends on whether you rent out excavators or sell apartments, so it is written once. Written again is the subject matter: the objects of the industry and the rules for moving between stages. A rental business has a unit of equipment, a booking and a dispatch. A property agency has a property, a listing and a deal. The words look alike, the rules inside do not, and it is the rules that cost time. ### How large is the share that has to be written from scratch? Larger than almost anyone expects, ourselves included at the start. Two figures, both measured. First: in the rental CRM we moved onto the shared base, 1650 of 1872 schema lines turned out to be industry-specific, that is 88%, and only 220 lines were the scaffolding the base provides. The second is harsher. When we built a second product in an adjacent industry, 7 of the 113 business concepts across the two products were shared. About 6%. Which gives a simple consequence for anyone choosing a vendor: the phrase we already have this, we just need to configure it describes the scaffolding, not your business. We wrote that second case up separately, because the number is inconvenient and deserves the detail. ### Why build a shared base at all if so little transfers? Because the saving does not come from where people look for it. Before a new product can deal with its own industry, it needs sign-ins, companies, roles and permissions, invoices wired to a real payment provider, incoming messages from customers, notifications, translations, a change log and encryption of personal data. The first time, that list takes months; it is identical for any industry; and any quiet mistake in it surfaces not in a demo but in a security audit. Built once, it lets the next product start where the business actually begins. The industry rules would have differed anyway, and the winnings were never there. ### Why publish a ledger of statuses instead of a product count? Because a product count is a marketing number and a ledger of statuses is an engineering one, and only the first can be wrong in a way nobody notices. The sentence eighteen products on a shared base will outlive any amount of drift underneath it: it stays true as long as anything at all is running. The line one in production, two in pilot, none in full conformance creates work, because every status is something somebody has to move. The ledger is kept by hand and updated on every release of the specification precisely so it does not quietly slide back into the first phrasing. The same reasoning keeps the specification repository free of executable code: a specification that ships a runtime stops being checkable against anything. --- # Browser-Agent-Ready SaaS: Making Your App Usable by Operator and Claude URL: https://inite.ai/en/blog/browser-agent-ready-saas Date: 2026-05-11 Author: Olga Fedotova Category: Agentic Engineering Tags: Browser Agents, Operator, ChatGPT Agent, Computer Use, Accessibility ## Direct Answer A browser-agent-ready SaaS is one where Operator, ChatGPT Agent, and Claude Computer Use can complete primary user flows (login, search, fill form, check out, retrieve a result) without human intervention. The five requirements are: (1) stable selectors that survive a redeploy, (2) form fields with semantic name/label/autocomplete attributes, (3) no Cloudflare-Turnstile / hCaptcha walls on read-only flows, (4) clear error states the agent can recover from, (5) a llms.txt-style agent manifest pointing at the action API. Sites that fail these requirements get abandoned by the agent in 60-90 seconds. ## Key Facts - OpenAI's Computer-Using Agent, the model behind Operator, scores 58.1% on WebArena and 38.1% on OSWorld — the reported human baseline on WebArena is about 78%. - Claude 4.5 Computer Use succeeds 87.4% on the OSWorld benchmark when forms have autocomplete attributes; 52.1% when they don't. - A challenge an agent cannot solve ends the session: Cloudflare's own guidance is to gate on risk rather than blanket-challenge, because an identified agent and a scraper are not the same visitor. - Pages with stable `data-testid` or `aria-label` attributes have 3.4x higher agent task-completion than identical pages relying on CSS class hashes. - By April 2026, 14% of SaaS sign-ups on tracked sites came from agent-driven sessions (up from 0.3% in April 2025). In 2026 your customer may not be the person who clicked your ad. It may be the agent they delegated the task to. OpenAI's Operator booked 1.2 million hotel rooms in Q1 2026. Claude Computer Use closes B2B SaaS trials. ChatGPT Agent fills out government forms. Browser Use runs at the bottom of every solo founder's automation stack. Each one of these is a vision-driven LLM that opens your site in a real Chromium browser, looks at the screen, decides what to click, and tries again on failure. They succeed when the page is semantically readable; they abandon when it isn't. The gap between "agent-friendly" and "agent-hostile" is now the same gap that mattered in 2010 for mobile and in 2018 for screen readers: not a luxury, a tier of customer. This is the 2026 audit checklist. ## How a browser agent sees your page Three perception modes, depending on the agent: 1. **Pure vision** (raw Operator, Browser Use's default): the agent takes a screenshot and asks the model "where should I click to do X?" Coordinates are the primary key. 2. **Vision + accessibility tree** (Claude Computer Use, ChatGPT Agent): the agent sees both the pixels AND the parsed ARIA/role/label tree. Far higher reliability because the model can target by name. 3. **DOM tap** (newer Operator builds, Browser Use's `dom_mode`): the agent reads the rendered DOM, extracts an enriched representation (selector + role + bbox + ancestors), and decides actions on structured data. You don't get to pick which mode your visitors use. So you design for mode 3 (the most demanding) and modes 1-2 inherit the benefit. ## The five requirements ### 1. Stable selectors Every interactive element you care about gets a `data-testid` or `data-agent-action` attribute. The value survives redeploys, brand refreshes, and tailwind upgrades. Examples that work: ```html View cart (3) ``` Examples that fail: ```html
``` If your team uses CSS-in-JS that emits hashed class names, the deploy-N selector and the deploy-N+1 selector are different strings. The agent's playbook (written by some upstream LLM with a 7-day-old training cut) breaks on every redeploy. `data-testid` is invariant by convention. ### 2. Semantic form attributes Every input gets at minimum `name`, `id`, `type`, `autocomplete`, and an associated `