For stalled or failed Microsoft Dynamics 365 Customer Engagement and Power Platform implementations. We run a forensic audit to diagnose the root causes, stabilize the environment, and deliver a working solution, including Field Service and its Resource Scheduling Optimization. Where it is already an emergency, the first 72 hours are set out step by step in the emergency rescue and stabilization process below, from environment lockdown to the vendor takeover of the keys. If you are planning a Customer Service rollout rather than repairing one, the Customer Service and Omnichannel implementation guide sets out the decisions, the deployment sequence, and the routing, knowledge, reporting, and adoption practice we would run it with, and the standard implementation timeline answers how long that takes, week by week, in eight to twelve weeks from Armenia. Where what stalled is a data migration rather than a build, the stalled migration rescue and takeover sets out the six week sequence we run, from finding the real deadline to a rehearsed cutover. If the money has already gone and a rescue cannot be bought yet, emergency stabilization with zero budget gives the six actions that cost nothing and the way to turn what they find into funding for the repair. If you are in Armenia, our local advantage and the 72 hour emergency takeover process for Armenia set out what a Yerevan based recovery looks like on the ground.
This is Solzet's specialized service for taking over and rescuing stalled or failed Microsoft Dynamics 365 Customer Engagement (CE) and Power Platform projects. We explain our forensic audit process to diagnose the root causes, our stabilization plan, and how we deliver a working solution. If your implementation is over budget, behind schedule, or your team has lost confidence in it, we can assume control and get it back on track. Solzet is a Microsoft Dynamics 365 Customer Engagement and Power Platform consultancy headquartered in Yerevan, Armenia, and picking up work that another partner or an internal team could not finish, including Field Service and its Resource Scheduling Optimization, is a routine part of what we do, either directly or on a white-label basis for other Microsoft partners.
A takeover is not a restart. In most cases the goal is to save the investment already made, keep the parts of the build that are sound, and fix only what is actually broken. We start by finding out what is really wrong, stabilize the environment so the bleeding stops, and then deliver against a plan you can see. Where a review shows the current design cannot be salvaged, we tell you plainly and scope the smallest rebuild that gets you to a working solution. The first six weeks are a fixed scope reset phase with dated milestones: an environment audit report inside two weeks, and a first working module in front of users inside four.
The project is over budget or past its deadline with no end in sight
Costs and timelines have overrun and every new estimate slips again. When the plan no longer predicts the finish, it usually means the underlying problems have never been diagnosed, only worked around.
Your original partner has stalled, gone quiet, or walked away
The implementation partner is no longer responsive, has run out of the skills the project needs, or the relationship has broken down. You are left holding an unfinished system and no clear owner.
Users have lost confidence and gone back to spreadsheets
The system is live in name only. People route around it, keep their real data in spreadsheets and email, and adoption is quietly collapsing because the build does not fit how they work.
The environment is unstable, slow, or full of errors
Flows fail silently, forms throw errors, integrations drop data, and no one is sure which customizations are safe to touch. Every change risks breaking something else.
Field Service scheduling is not producing usable results
Resource Scheduling Optimization runs but returns schedules the dispatch team cannot trust, or the schedule board is still worked by hand because the automated engine sends the wrong technician to the wrong job.
Nobody can explain how the solution actually works
There is little or no documentation, key knowledge left with the people who built it, and the current team cannot safely extend or support what exists. The system has become a black box.
Emergency Rescue & Stabilization Process
Not every troubled project is an emergency. The signals below are the narrower ones that say the situation cannot wait for a discovery phase, and what follows is what we actually do in the first seventy two hours when they are present: lock the environment down, take the keys back from the outgoing vendor, stop the damage, diagnose under pressure, and put one honest line of communication back in place before anyone talks about a recovery plan.
It is set out three ways, because a team in crisis reads differently from a team planning: the escalation signals that say this is an emergency, the stabilization steps in full, and then the same seventy two hours as a flowchart with the gates that change what happens next, followed by the vendor takeover checklist of what control moves and when.
The signals that make this an emergency
Any one of these on its own is reason to intervene this week rather than next quarter. Two or more together usually mean the project will not recover on its current path.
Go live has been missed more than once
A slipped date is normal. A second and third date set without anything changing about the reason the first one was missed is not, because it means nobody has diagnosed the cause and the plan is now a hope. By the third date the organization has usually stopped believing any of them, which is its own problem: the business has stopped preparing for a go live it no longer expects.
The business is being disrupted right now
Invoices are not going out, cases are not reaching the agents who own them, quotes are being rebuilt by hand, or dispatch has gone back to a whiteboard because the schedule board cannot be trusted. Once a broken implementation is costing revenue or service every day it runs, the cost of waiting for a proper discovery phase is higher than the cost of the discovery.
Data is being damaged, not just missing
A flow is overwriting records, an integration is writing duplicates faster than anyone can clean them, or a process is deleting rows nobody has confirmed are recoverable. This is the one signal that should never wait, because platform restore windows are finite and the decision about recovery has to be made while a restore point still exists.
The team is burning out and people are leaving
Weekend firefighting has become the normal way the system runs, the same two people are the only ones who can fix anything, and one of them has resigned or is about to. Burnout on a stalled project is not only a human cost. Every departure takes undocumented knowledge with it and makes the eventual rescue longer and more expensive.
Changes are still going straight into production
Fixes are being typed into the live environment because there is no working path from development to test to production, so every attempt to make things better is also a new risk. When the only way to fix the system is to endanger it, the environment has to be frozen before anything else is decided.
The written status and what users say no longer match
Reports say green while the people using the system describe something else, and no one will put the actual state of the build in writing. That gap is usually the last signal before a project fails outright, because it means decisions are being made on a picture that stopped being true some time ago.
The stabilization process, step by step
Stabilization is a separate thing from recovery and it comes first. Its only job is to stop the situation getting worse and to establish what is true, so that the plan built afterwards is built on facts rather than on the account of the project everyone has been given so far.
1
Hours 0 to 24
Environment lockdown
The first action is not analysis, it is stopping uncontrolled change. We freeze deployments into production, cut the System Customizer and System Administrator roles back to the people who genuinely need them, and stop edits being made directly in the live environment. We take a copy of production into a sandbox so the state that caused the crisis is preserved and can be examined without touching what users depend on. Flows, jobs, and automatic record creation rules that are actively damaging data are turned off rather than rewritten, and we confirm what restore points exist and how long the environment keeps them, because that window closes whether or not anyone is watching it. Nothing is deleted during lockdown. The purpose is to hold the system still.
2
Hours 0 to 72
Emergency triage of what is hurting people today
Alongside the lockdown we fix the failures blocking daily work: the form that will not save, the flow overwriting records, the integration dropping rows, the schedule sending crews to the wrong side of the city. Each one is reproduced in the sandbox copy, fixed there, and promoted through a controlled path rather than typed into production. This runs before the audit is finished on purpose. A team that has already been told twice to wait for a plan will not wait again, and visible relief in the first days is what buys the rescue the credibility to continue.
3
Days 1 to 3
Forensic audit under emergency conditions
An emergency audit answers three questions first: what is actively causing harm, what is going to fail next, and what must not be touched until the full picture exists. We read the environment itself, the solution layers, where business logic actually lives, the Dataverse model, security roles, integrations, and the accumulated failures in flow and plug-in run history, rather than the documentation about it. The full assessment method, including the three root cause families we test for and a worked remediation, is set out in our rescue services guide, and the same inspection run outside a crisis is our health check and technical audit. The emergency version is the same discipline, prioritized by what is bleeding.
4
Days 1 to 3
The communication plan
A project in crisis is usually running three contradictory stories at once, and that is as damaging as any technical fault. We fix a single named decision maker, one channel of record where decisions are written down, and a daily written brief at a fixed time saying what was found, what changed, and what is next. Users are told plainly what is frozen, what is being fixed, and what to do in the meantime, because silence is what sends people back to spreadsheets for good. The sponsor gets the same brief as everyone else, so there is no separate version of reality to reconcile later.
5
End of week 1
The stabilization statement and the decision point
The intervention closes with a short written statement: what was locked down, what was fixed, what is still dangerous, what data damage is recoverable and what is not, and what it will take to move from stable to working. That statement is yours whatever you decide next, and it is written to be actionable by your own team or another partner. If you continue with us, it becomes the front door to the fixed scope reset phase below, which carries a fixed price, a fixed end date, and a named deliverable at every milestone.
The same seventy two hours drawn as the path rather than described as a promise. Read it top to bottom. Each box is either something we do or a question whose answer changes what happens next, and the two branches under a question are what we do in either case. It is written this way so you can find where your own project would enter the process and see what would happen to it in the following hour.
1Hour 0Step
The call, and the clock starting
An emergency takeover needs two people before it needs anything else: someone on your side who can authorize access to the environment, and a named person on ours who owns the intervention from this hour until the stabilization statement is written. We start on read access where full administration rights take longer to arrange, because waiting for permissions is the most common way an emergency loses its first day. The clock below starts when access is granted rather than when a contract is signed.
?Hour 0 to 2Decision point
Is anything damaging data right now?
This is the only question asked before lockdown, because it is the only failure whose cost grows every hour and whose remedy has an expiry date. A flow overwriting records, an integration writing duplicates faster than they can be merged, or a process deleting rows nobody has confirmed are recoverable all belong here.
If yes
The thing doing the writing is turned off before anything else is discussed, and we confirm what restore points exist and how long the environment keeps them while a restore point still exists to use. That window closes whether or not anyone is watching it, so the recovery decision is made inside it or not at all.
If no
We go straight to lockdown. Nothing is deleted, nothing is rewritten, and no opinion is offered about the quality of the build until the environment has stopped moving.
3Hour 0 to 8Step
Environment lockdown
Deployments into production are frozen, System Customizer and System Administrator are cut back to the people who genuinely need them, direct edits in the live environment stop, and production is copied into a sandbox so the failing state is preserved. Nothing is deleted at this gate. Its only job is to hold the system still, so that everything measured afterwards is measuring one system rather than a moving one.
4Hour 2 to 12Step
Vendor takeover: the keys change hands
A vendor takeover is an administrative act before it is a technical one, and it is the step that stalls most often because nobody on the client side realizes it can simply be insisted on. A Dynamics 365 environment is tenant property. Tenant administration, environment ownership, the identities the integrations authenticate as, and the repositories holding the solution source can all be recovered by whoever owns the tenant, regardless of who built the system or how the relationship ended. What moves, and in what order, is set out below the flowchart.
5Hour 4 to 24Step
Emergency triage of what is blocking work today
This runs alongside the audit rather than after it: the form that will not save, the flow overwriting records, the integration dropping rows, the schedule sending crews to the wrong side of the city. A team that has already been told twice to wait for a plan will not wait a third time, and visible relief on the first day is what buys the rescue the credibility to finish.
?Hour 8 to 24Decision point
Does the failure reproduce in the sandbox copy?
Every emergency fix is reproduced before it is written. This one gate is most of what separates a stabilization from more of whatever caused the crisis.
If yes
It is fixed in the sandbox and promoted into production through a controlled path, one change at a time, so that a new problem points at a single change rather than at a batch of them.
If no
It is not a code fault. It is data, permissions, or environment specific configuration, and it goes onto the audit list instead of being chased in production. Typing a fix into a live environment for a failure nobody can reproduce is how the previous team arrived here.
7Hour 12 to 72Step
Forensic audit under emergency conditions
Three questions in priority order: what is actively causing harm, what is going to fail next, and what must not be touched until the full picture exists. We read the environment itself rather than the documentation about it, and the audit is scoped by what is bleeding rather than by what would be interesting to know. The full method lives on the rescue services guide and the health check page rather than being repeated in a crisis.
8Running from hour 0Runs in parallel
One decision maker, one channel, one daily brief
The communication plan is not something that arrives at the end of the week. It runs in parallel with every box above, from the first hour: one named decision maker, one channel of record where decisions are written down, and a written brief at a fixed time each day saying what was found, what changed, and what happens next. Users are told plainly what is frozen and what to do in the meantime, because silence is what sends people back to spreadsheets permanently.
?Hour 72Decision point
Is the environment stable enough to plan from?
Stable is not the same as fixed. It means nothing is getting worse on its own, the damage is bounded and written down, and a change can be made and released without guesswork.
If yes
We write the stabilization statement and the fixed scope reset phase can begin, with a fixed price, a fixed end date, and a named deliverable at every milestone.
If no
Stabilization continues into a second week with the scope narrowed to whatever is still bleeding, and we say so in the daily brief on the day we know rather than at the end of it. A reset planned on an environment that is still moving is a plan that gets rewritten.
10End of week 1Outcome
The stabilization statement and the decision point
One page, written the same week: what was locked down, what was fixed, what is still dangerous, what data damage is recoverable and what is not, and what it will take to move from stable to working. The statement is yours whether you continue with us, hand it to another partner, or carry on with your own team.
Vendor takeover: what changes hands, and in what order
Box four of the flowchart is the one that decides how fast the rest of it can run, and it is administrative rather than technical. This is what a vendor takeover actually consists of on a Dynamics 365 and Power Platform project, on the clock we work it to. None of it requires the outgoing party to cooperate, which is the part worth repeating to anyone who has been told the keys are gone: the tenant, the environments, and the data are yours.
1
Hour 0 to 4
Tenant administration
Global Administrator on the Microsoft 365 tenant is where every other permission is granted from, so it moves first. On a broken engagement the outgoing partner often holds it, and clients regularly read that as the keys being gone. They are not. The tenant belongs to the organization paying for it, and administration can be recovered through the tenant owner without the outgoing party agreeing to anything.
2
Hour 2 to 8
Power Platform and Dataverse administration
Power Platform Administrator, environment ownership, and Dataverse System Administrator for the people who will now own the system, across production and every sandbox and developer environment attached to the project. Environments nobody mentioned in the handover conversation are a routine find at this step, and they matter, because one of them is usually where the last set of changes was really made.
3
Hour 4 to 24
The identities your integrations run as
Every interface authenticates as something: an application user backed by an app registration, a service principal, or, more often than anyone expects, a named person who has already left the project. We record what each integration runs as, take ownership of the app registrations and their secrets, and note every certificate and secret expiry date. A takeover that skips this ends with an interface failing silently a few weeks later for a reason nobody connects back to the handover.
4
Hour 8 to 24
Source, pipelines, and the deployment path
The repositories the solution source actually lives in, the pipelines that deploy it, and the service connections those pipelines authenticate through. Where the source exists nowhere outside the environment, that is a finding rather than an obstacle: exporting the unpacked solution into a repository you own becomes the first item of the reset, because until it exists there is no version of the system that survives the environment it runs in.
5
Day 2 to 3
Licensing and the commercial edges
Who holds the licensing relationship, what the contract says about ownership of the custom code, and whether anything in production was supplied by the outgoing party as intellectual property they intend to keep. This is slower than the technical steps and it is deliberately not on the critical path, but it has to be asked in the first week. Discovering in month two that a plug-in the business depends on is not yours is a problem with no cheap answer.
6
Whenever it can be had
The last hour with the outgoing team
The cheapest hour in the entire rescue, and the one most often skipped because the relationship has soured. It answers what no audit can recover: which customization was a deliberate design decision and which was a workaround for something else, what was promised verbally and to whom, and which integration nobody dares touch and why. We ask for it in writing, keep it short, and go in with a prepared list of questions rather than a conversation.
Where the outgoing team is still available, the handover is a relationship as well as a permission change, and the mechanics of that, including how knowledge is extracted before the last people leave, are set out in the takeover process further down this page. Where they have already gone, the audit reconstructs the design from the environment itself, which is slower and is ordinary work rather than a blocker. If nothing here is a crisis and you simply want the build continued by a team that knows the platform, that is ordinary Dynamics 365 consulting and development rather than a rescue, and it starts from the same read of the environment as a health check and technical audit.
Emergency stabilization compared with a standard consulting engagement
Most rescue proposals are an ordinary consulting engagement with the word emergency on the cover: the same discovery phase, the same time and materials, the same report at the end of it. The difference is not the intent, it is what happens in the first week.
In an emergency
Standard consultant approach
Solzet emergency stabilization
The first move
Discovery workshops scheduled to start in a few weeks, with the environment left exactly as it is in the meantime
Environment lockdown on day one, then triage, so uncontrolled change stops before anything is analyzed
Changes to production
Fixes continue to be made in the live environment while the assessment runs, adding new layers to the problem being assessed
Production is frozen, work happens in a sandbox copy, and changes are promoted through a controlled path
What the diagnosis reads
Documentation, project plans, and interviews with the people who built it
The environment itself: solution layers, code, flows, security roles, integrations, and run history, compared against how the business works
When you first see something
A findings deck at the end of a discovery phase, typically weeks away
A daily written brief from day one, emergency fixes released inside the first week, and a written stabilization statement by the end of it
Commercials
Time and materials that expands with the scope of what is found
A fixed scope reset phase with a fixed price, a fixed end date, and a named deliverable at every milestone
Who does the work
An assessment team writes the report and hands it to a delivery team who were not there for the diagnosis
The people who diagnose it are the people who fix it, so nothing is lost in a handover
How it ends
A proposal for a long programme, with the report treated as the vendor property
A go or stop decision that is yours, with the audit, the debt inventory, and the handover pack written to be usable by any partner or your own team
Stabilization is where a takeover starts, not where it ends. Once the environment is still and the facts are written down, the work below is what gets the implementation to a solution your team can actually own.
Emergency Stabilization with Zero Budget
Most rescue advice assumes the reader can buy a rescue. Often they cannot. The money went into the implementation that failed, the partner who spent it has stopped answering, and the people left are being asked to keep the business running on a system nobody trusts. If that is where you are, everything in the first half of this section is free, needs no purchase order, and does not need us. Do it this week rather than waiting for anyone to quote for it.
The second half is the part free crisis advice usually leaves out. Stabilizing a broken system without money keeps it from getting worse and fixes nothing, so the real question is how the work gets paid for. Each action below therefore produces a specific piece of evidence, and the funding case that follows is built out of exactly those pieces. Nothing here is a sales route: it works whether the money ends up with us, with the partner who built the system, or with a hire on your own payroll.
What to do this week, for nothing
In this order. The sequence matters more than the speed, because the first two actions protect your ability to do the rest, and one of them has an expiry date you do not control.
1
Hour 1Free. Thirty minutes in the admin centers.
Secure administrative access before anything else
Everything else on this list assumes you can still get into your own system, and that assumption expires. Confirm that at least two people inside your organization hold Global Administrator and Power Platform Administrator, that someone other than the outgoing partner holds System Administrator in each environment, and that you know the account and the credentials behind every integration and every connection reference your flows run on. The tenant, the environments, and the data are yours regardless of who built the system and regardless of how the relationship ended, so none of this requires the other side to agree to anything. The mechanics of recovering the keys where the outgoing party has stopped answering are set out in the vendor takeover checklist above; do the first two boxes of it yourself today.
What this gives the funding case: A written list of who can still change your system. If the only administrator is a contractor on notice or a partner who has stopped replying, that is a dated risk rather than an opinion, and dated risks are what get budget released.
2
Hours 1 to 4Free. One written message and one manual backup.
Freeze change and find the date your restore window closes
Send one message saying that no new configuration, no new flows, and no new deployments go into production until further notice, and name the single person who can lift the freeze. Then make it real by cutting the System Customizer and System Administrator roles back to the few people who genuinely need them. While you are in the Power Platform admin center, take a manual backup of the production environment and read the retention period your own tenant actually shows, then write down the date the oldest usable restore point expires. Do not restore anything yet. The point today is only to be sure that a point to go back to still exists next week.
What this gives the funding case: A date that is not negotiable and is not yours. Budget moves against deadlines set by something other than the person asking, and the day your restore window closes is the one deadline in this crisis nobody can argue with.
3
Day 1Free. One day of one person who knows the business.
Document the customizations while somebody still remembers them
Open Solutions in the maker portal and export the unmanaged solutions holding your customizations, then export a copy of any Power Automate flows that live outside a solution, because those are the ones that vanish quietly when a licence lapses or a person leaves. Alongside the files, write a plain inventory: for each module, what was built, who asked for it, whether it currently works, and who last touched it. List every integration and the account it authenticates as. Ask the outgoing partner in writing, now, for the source code of any plug-ins or custom controls, and check whether your contract already says it belongs to you. Export the records the business cannot lose at the same time, dated and stored outside the tenant.
What this gives the funding case: A written inventory of the build. This is the single largest discount available on any eventual quote, because every hour a rescue partner spends reconstructing what was built is an hour billed at consulting rates. An inventory your own team writes for nothing is work no one has to pay for twice.
4
Days 1 to 2Free. An afternoon reading run history.
Disable the automations that are actively damaging data
Damage to data is a different problem from a system that merely does not work, because it compounds every day and it is the hardest thing to undo later. Go through everything that writes on its own: Power Automate flows, classic workflows, automatic record creation rules, bulk delete jobs, and any integration writing into Dataverse. Read the run history for the flows failing repeatedly, and then for the more dangerous case, the flow that succeeds while doing the wrong thing. Switch off anything overwriting, duplicating, or deleting records. Turn off rather than delete, and keep a plain list of what you disabled, when, and why, so it can be put back deliberately rather than rediscovered by accident. Where you cannot tell whether something is doing harm, ask whether anyone would notice if it stopped for a week.
What this gives the funding case: The list of what had to be switched off to stop the bleeding. Eleven disabled automations on one page is a more persuasive paper than any adjective, because it says without argument that the system is not currently doing the job it was bought to do.
5
Days 2 to 5Free to set up. Expensive to run, which is the point.
Establish one manual fallback per broken process, and count what it costs
While the system is frozen people still have to work, and the real danger is that every team invents a different workaround. Pick one per process and mandate it. One shared mailbox or one queue worked by a named person on a rota, rather than mail sitting in individual inboxes. One spreadsheet, in one place, with column headings copied exactly from the fields in Dynamics 365, so what is captured by hand can be loaded back later instead of retyped. One rule that every row carries a date, a person, and the record it belongs to. Tell users plainly what is frozen, what to do instead, and when they will hear next, because silence is what sends a team back to spreadsheets permanently. Then do the part almost everyone skips: have each team log the hours the fallback costs them, every week, from the first week.
What this gives the funding case: Hours per week, by team, with dates on them. This is the number that turns the funding conversation from an argument about a failed project into arithmetic, and it is the only item on this list you cannot reconstruct later if you did not start counting.
6
End of week 1Free. An hour.
Write the state of things down in one page while it is still fresh
One dated page: what is frozen and who can lift it, what was disabled and why, which processes are being run by hand, what data is known to be wrong, when the restore window closes, and who holds the keys. Circulate it to the sponsor and to the people doing the manual work, and update it weekly. A project in crisis is normally running three contradictory accounts of itself at once, and that costs as much as any technical fault. This page ends that, and it is the document any partner, any new hire, or your own team six weeks from now starts from. It is also the document the funding decision is made on, because a sponsor will not release money against a description of a mess, only against one page with dates, numbers, and a named next step.
What this gives the funding case: The funding paper itself, most of it already written. Everything below is about what to put around it.
The longer form of this playbook, with the exports and the admin center steps written out click by click, and with the five conditions under which carrying on alone stops being prudent and starts being expensive, is in Dynamics 365 rescue services: how to fix a failed or stalled implementation. Read those limits before you attempt any repair beyond the six actions above, because damaged data corrected by hand and a production restore taken under time pressure are the two places where well meant DIY does more harm than the original failure. For the wider question of which work is configuration your own team can genuinely carry and which sits behind a wall that needs a developer, see when to hire a Dynamics 365 consultant versus a DIY implementation.
Quantify the business risk, because that is what gets funded
A broken system is not funded because it is broken. It is funded when somebody can see what it is costing per week and what becomes impossible after a specific date. Every category below can be measured with numbers you already hold, and none of it needs an outside party to produce. There are no benchmarks in this table on purpose: the figures that persuade your own board are your own.
The risk
How to put a number on it
Where the number comes from
The manual fallback
Hours a week across every team affected, times a loaded hourly cost, times the number of weeks before anything changes
The tally started in the fifth action above. Use loaded cost, salary plus employer cost, rather than salary, because that is the figure finance will recognize
Work the system is no longer doing
The volume that used to pass through the system in a normal week against what is passing through it now, expressed as quotes delayed, invoices late by days, cases past their promised response, or jobs dispatched by phone
Your own operational reporting for the last month before the crisis, compared against last week. Where the reports themselves have stopped working, count one week by hand
Data damage that is about to become permanent
The date the oldest usable restore point expires, plus the count of records already known to be wrong. After that date, correcting them stops being a restore and becomes a manual project priced by the record
The backup retention shown for your own environment in the Power Platform admin center, and record counts from the exports taken in the third action above
Money still being spent on a system nobody is using
Licences billed every month for users who are working around the system, plus premium capacity and any managed service still being invoiced
Your Microsoft 365 admin center billing against actual usage. This is a current loss rather than a sunk one, which matters: sunk cost never justifies more spending, but a live monthly bill for an unused system does
Knowledge that is one resignation from leaving
Name the two or three people who could leave and write down what stops working the day they do. Where notice has already been given, put that date on the paper
The administrator list from the first action and the inventory from the third. A build that only one person understands is a risk with a date on it, not a staffing preference
Compliance, audit, and contractual exposure
Which records carrying a legal, contractual, or audit obligation are now being held in spreadsheets outside the system, and where an audit trail has stopped being written at all
Your own retention and reporting obligations, tested against the manual fallback you just mandated. This is normally the row that moves a board when the labour arithmetic alone does not
The same thing as arithmetic, with your numbers in place of these
Eleven people spending six hours a week each on the manual fallback is sixty six hours a week. At a loaded cost of thirty per hour in whatever currency you report in, that is 1,980 a week, or roughly 25,700 across a quarter, and it is the cost of the workaround alone. It counts nothing for the quotes going out late, nothing for the licences still being billed every month for people who are working in a spreadsheet, and nothing for the records that stop being restorable on a date already in the calendar. Set that against the price of a bounded assessment, which is a number one person can approve in a single meeting, and the case stops being a debate about a failed project and becomes a comparison between two figures. The numbers here are placeholders to show the shape of the calculation. Substitute your own and the shape holds.
Turning the numbers into approved money
Teams in this position rarely fail because the business refuses to spend. They fail because the request is written as a rescue of unknown size, with no date on it, and taken to the budget holder who has already paid for this project once. Five corrections.
1
State the loss as a rate, not a total
A total invites an argument about how it was calculated. A rate invites a decision, because it keeps running while the decision is not made. Put the labour, the delayed work, and the licence spend on one line as a cost per week, show the arithmetic underneath it in three rows, and say what it will have reached by the next board meeting if nothing changes. Round down and say that you rounded down. A conservative number that survives being challenged is worth more than an accurate one that does not.
2
Attach a deadline that did not come from you
Requests without dates are approved in the order of the dates attached to them, which means never. You already have real ones: the day your restore window closes, the day a source system stops being readable, the last day of a departing administrator, a licence renewal, a contractual reporting date, an audit. Name the earliest one, say plainly what becomes impossible or more expensive after it, and put it in the first three lines. This is the difference between a request that sits in a queue and a request that has to be answered this month.
3
Ask for the smallest thing that unblocks the next decision
The most common reason a rescue is not funded is that somebody asked for the rescue. A sponsor who will not approve a programme of unknown size will approve a bounded, fixed price read of the environment, because it ends in a decision rather than a commitment and its price cannot grow with what it finds. So the ask is an independent assessment, not a recovery, and the paper says exactly what comes back from it and when. Our version of that is a fixed price health check and technical audit whose deliverable is a prioritized remediation plan you own and can hand to any partner or run yourself, including the one who built the system.
4
Price three options, not one
One option asks for a yes or a no, and under pressure the safe answer is no. Three ask for a choice. Option one is doing nothing, and it is not free: its price is the weekly rate from the first step, carried to the end of the quarter. Option two is stabilization only, which stops the loss growing and leaves the build where it is. Option three is a full recovery to a working system. Give each one a number, a date, and what the business gets, and let the reader pick. Presenting the cost of inaction as a priced option rather than as a warning is the single change that most often turns this conversation.
5
Take it to the person carrying the loss, not the person holding the IT budget
The numbers you have assembled are operational: hours, delayed invoices, cases past their promise, jobs dispatched by phone. They belong to whoever runs that operation, and that is usually a service, sales, or operations director whose team is doing the work by hand every day, not the IT budget line that has already been spent once on this project and is being defended. Show them the one page, say what you have already done for nothing, and ask them to sponsor the smallest option. Saying plainly that the team has already frozen the environment, disabled what was doing harm, documented the build, and stood up the fallbacks matters more than it looks: it is the difference between asking to be rescued from a mess and asking to finish a job already half done.
When the smallest fundable step is the ask, this is what it looks like on our side: a fixed price Dynamics 365 health check and technical audit, agreed before we start, that does not move with what it finds and ends in a prioritized plan you own and can hand to anyone. Where the situation is past assessment and something has to be stopped this week, the emergency stabilization process above is the paid version of the first two actions on this list, and the vendor takeover checklist covers recovering the keys when the outgoing party has stopped cooperating. If the answer is that there is genuinely no money at all this quarter, do the six free actions, keep counting the hours, and come back with the arithmetic. We would rather tell you that on a page than on a call.
How our rescue and takeover process works
The same sequence every time: find out what is really wrong, make the system safe, agree a plan you can see, deliver the fixes in visible increments, and hand back a solution your team can own.
1
Forensic audit to diagnose the root causes
We start by examining the actual solution, not the story around it. We review the Dataverse data model, forms and views, business rules, plug-ins and custom code, Power Automate flows, security roles, integrations, and any PCF or canvas components, and we compare what was built against what the business needs. The output is a clear, written diagnosis: what is sound, what is broken, what is missing, and which problems are causing the visible symptoms. This is the evidence every later step depends on.
2
Stabilize the environment so the bleeding stops
Before adding anything new, we make the system safe to run and safe to change. That means fixing the failures that lose data or block daily work, containing risky customizations, tightening security roles, and getting environments, solutions, and source control into a state where a change can be made and released without guesswork. Stabilization buys the team room to breathe and stops the situation getting worse while the plan is built.
3
Agree a stabilization and recovery plan you can see
From the audit we build a prioritized plan: the specific fixes, the order to do them in, what we keep, what we rework, and what, if anything, needs rebuilding. Each item is tied to a business outcome so you can see why it matters and what done looks like. Where the honest answer is that part of the design cannot be salvaged, the plan says so and scopes the smallest rebuild that gets you to a working solution rather than a full restart.
4
Deliver the fixes and complete the build
We work through the plan in short, visible increments, correcting the model driven and canvas apps, repairing or rewriting flows and code, fixing integrations, and finishing the functionality the project was supposed to deliver. You see working software regularly rather than waiting for one distant go live, and each increment is tested against the real process it supports.
5
Harden, document, and hand back control
A rescue is only finished when your team can run the system without us. We document how the solution is built and why, put in place a clean solution and deployment approach, and transfer knowledge to your people. You end with a working, supportable Dynamics 365 and Power Platform solution and the documentation and confidence to own it, whether we stay on for ongoing support or step away.
Common Rescue Scenarios and Our Fixes
The process above is the same on every takeover. What actually differs is the failure pattern underneath it, and that is where most rescue advice stops: a three phase diagram of assess, stabilize, and optimize that is true of every project and useful on none of them. The four patterns below are the ones we meet most often on inherited Dynamics 365 Customer Engagement and Power Platform projects, and each one is set out the way we would work it: what it looks like from the inside, the specific checks that confirm it rather than assume it, the remediation in the order the platform forces it to happen, and what done looks like when it is finished.
The diagnostics are written so your own team could run them this week without us. That is deliberate. A rescue section is only worth more than a diagram if the person reading it can use it to find out what is wrong with their project before they have hired anybody.
Scope creep that turned a phase one into an open ended programme
What it looks like
The build never shipped because it never stopped growing. Phase one was a quarter and is now in its second year, the backlog is larger than it was at kickoff, and nobody can point at the document that says what the first release contains. The tell is not the size of the backlog, it is that new items enter it without displacing anything and without a name against the approval. Left alone this pattern does not end in a bad system, it ends in no system, because there is no definition of done that anyone can reach.
How we confirm it, step by step
1
Mark the backlog against the original signed scope
Take the scope that was actually agreed and go through the current backlog item by item, marking each one in scope, added later, or a reworded duplicate of something already delivered. On a project in this state the third category is routinely a quarter of the list, because a delivered item that failed to satisfy someone comes back with a new title instead of a defect. Counting it is what turns an argument about whether there is scope creep into a number the sponsor can look at.
2
Put a date and a name on every change request
For each added item, record when it was raised, when it was approved, and who approved it. What you are looking for is the month approvals stopped carrying a name, because that is when change control failed and everything after it entered the project by default rather than by decision. It is also the fairest evidence available, since it points at a broken process rather than at a person.
3
Trace each added item to a decision the business makes
Every requirement should be traceable to a decision somebody takes or an action somebody performs. Ask, for each added item, which decision it supports and who takes it. Items that cannot be traced are usually a preference, a report someone wanted once, or a field added so a spreadsheet could be retired without anyone checking whether it still is. They are the cheapest scope to cut because no process depends on them.
4
Find the scope that was built and never used
Read the environment for work that was delivered into a vacuum: model driven apps with no security role assigned to any user, tables created a year ago holding a handful of rows, flows turned off after the first week, forms nobody is assigned to, business process flows with no active instances. This is scope that has already been paid for and is producing nothing, and it is also maintenance and testing cost on every release. Cutting it is free.
5
Read the acceptance criteria on what was called done
Pull ten delivered items at random and read how done was defined. Where the criterion is a restatement of the requirement rather than something a user could demonstrate, the item will be reopened later and belongs back in the count. This check is what stops a re-baseline from being built on a delivered list that is not really delivered.
How we fix it, phase by phase
Week 1
Freeze the backlog and stop the intake
Nothing new enters the backlog until a scope is signed, and that is written down and told to everyone who raises requests rather than quietly enforced. Anything urgent that arrives during the freeze goes on a single parked list with a date, which matters because a freeze that has no visible place to put an urgent request is a freeze people go around. The freeze is not a judgement about the requests. It is the only way to count what exists.
Week 2
Re-baseline to a minimum usable scope
Every surviving item is marked ship now, ship later, or drop, and the result is the one page Reset Scope described below, with acceptance criteria written in language a user would recognize. Ship now is deliberately narrow: one module a team can work end to end. The drop list is written down and shown to the people who asked for those items rather than silently deleted, because scope that is dropped without being seen to be dropped comes back within a month and the re-baseline was wasted.
Week 3 onward
Reopen change control with a displacement rule
Change control restarts with one named decision maker, a written request that states the decision the item supports, and a rule that any new item of a given size displaces something of equal size from the current release rather than adding to it. The weekly thirty minute checkpoint is where displacement is decided, in front of the same people every week. This is the part that makes the fix hold, because a re-baseline without a displacement rule simply restarts the same growth from a lower number.
What done looks like
The release contains a fixed, written list of items with demonstrable acceptance criteria, the drop list is visible to the people who asked for those items, and every change since the re-baseline carries a name, a date, and the thing it displaced.
A data migration that brought the wrong data across
What it looks like
The records are in Dataverse and the numbers are wrong. Reports do not match the system they came from, pipeline and case ageing look impossible, duplicates multiply faster than the team can merge them, and users have stopped trusting anything the system tells them. Migration damage is the failure pattern with the hardest deadline in a rescue, because a source system is often switched off or read only on a date somebody already agreed, and after that date some of what is missing stops being recoverable.
How we confirm it, step by step
1
Reconcile counts twice, whole table and business subset
Compare row counts per table between the source and Dataverse, then compare again on the subset the business actually cares about, such as active customers, open cases, or opportunities closed in the last two years. Whole table counts frequently match while the subset does not, which is the signature of a filter applied at extract that nobody recorded. A count that matches on the total and fails on the subset is more informative than either number alone.
2
Plot created on and look for a spike on load day
Group records by created date and look at the distribution. If the dates cluster on the migration date instead of spreading across the years the source covers, the load did not set the created on override, and that cannot be corrected afterwards on those rows: it has to be set when a record is created. Every ageing report, SLA calculation, and trend line built on that table is wrong until it is reloaded, so this check decides scope before any other correction is planned.
3
Check who owns the records
Count records by owner. Where a table is almost entirely owned by the migration account or by whichever administrator ran the load, the security model has silently stopped working: user and team scoped roles no longer show people their own data, assignment rules behave oddly, and anything driven by ownership reports against the wrong person. This is common because ownership has to be mapped and set during the load rather than fixed with a bulk edit later.
4
Check state, status, and currency on the rows that carry money
Closed records that arrived open inflate the pipeline and the case backlog and pull open items into views and automation that should never have seen them. Separately, look at every money column: records loaded without a transaction currency fall back to the base currency, so the numbers are arithmetically fine and factually wrong. Both are correctable in place, which is why they are worth separating from the damage that is not.
5
Read the load logs for lookups that never resolved
Go back to the migration job errors or the staging tables and count the rows where a lookup failed to match and was written as empty or dropped into a text column. Orphaned children are the most common form of quiet migration damage: the activity that lost its regarding record, the contact with no parent account, the case attached to nothing. They do not show as errors in the application, they show as work that disappeared.
6
Measure duplication with the rules rather than by eye
Run duplicate detection jobs across the core tables and read the match counts instead of sampling the grid. A rescue needs the size of the duplication problem, not proof that it exists, because the number decides whether the answer is a merge exercise, a reload against an alternate key, or a change to how records are created before anything is cleaned at all.
7
Confirm whether history and attachments came at all
Notes, attachments, and activity history are usually a separate workstream from the parent records and are the part most often deferred and then forgotten. Check the counts and, for files, the total bytes, because a partial file migration reports a healthy row count while the documents themselves are truncated or absent.
How we fix it, phase by phase
Days 1 to 3
Stop the writing and preserve what can still be read
If a delta or incremental load is still running on a schedule, turn it off first. A migration that is still writing will undo corrections faster than they can be made, and it is the single most common reason a data cleanup runs twice. Then preserve evidence: keep the staging data, take a copy of production, and confirm how long the environment keeps restore points. Finally establish whether the source system is still readable and until when, because that date, not our plan, sets the deadline for anything that has to be re-extracted.
Week 1
Classify every finding as reload, correct in place, or lost
Reload applies where the source is still authoritative and users have not yet edited the migrated rows, and it is the right answer for anything the created on override missed. Correct in place applies where users have worked on the data since go live, because a reload there would destroy real work that exists nowhere else. Lost is the honest third category: data that is in neither system and has to be accepted or recaptured manually. Writing that third list down, and telling the business what is on it, is what stops a rescue from being blamed later for damage it inherited.
Weeks 2 to 4
Correct in dependency order, never in parallel
Corrections run in the order the platform requires and one pass at a time: users and teams, then ownership, then currencies and reference and choice data, then lookups and parent records, then state and status, then the transaction rows themselves. Correcting rows before the records they point at exist simply produces a second pass. Every pass runs first in a sandbox copied from production, with counts taken before and after, and automation, duplicate detection, and auditing are disabled for the duration of a bulk pass and re-enabled from a written checklist rather than from memory.
Weeks 4 to 6
Prove it, then make the same damage impossible
The correction closes with a reconciliation per table that the business signs rather than a statement that the data is fixed: counts, the subset counts, created date distribution, ownership spread, duplicate match counts, and file bytes. Prevention is configuration and it is quick: alternate keys on the tables that receive integrations, duplicate detection rules published and running on a schedule, and required mappings on the interfaces that create records. Without that last phase the same duplicates return through the same door within a quarter.
What done looks like
Counts reconcile at both the table and business subset level, created dates span the period the source covers, ownership matches the real users, closed records are closed, and the duplicate rate is measured and falling rather than argued about.
Customization debt that makes every change break something else
What it looks like
The system works until somebody touches it. Releases are avoided, small changes take weeks, the same two people are the only ones who dare deploy, and an update from Microsoft is treated as a threat rather than a routine event. Underneath it is almost always the same shape: business logic scattered across several places that all fire on the same record, unmanaged layers sitting on top of managed ones, and no reliable path from development to production. This is the pattern where the temptation to rebuild is strongest and usually wrong, because most of the debt is removable without touching what the business runs on.
How we confirm it, step by step
1
Read the solution layers and find what is winning
For the tables and forms that matter, open the layers view and look at the stack above the base component. An unmanaged layer over a managed one wins permanently and is the specific reason an update or a partner fix appears to have been installed and to have done nothing. Record which components have unmanaged layers on them, because that list is the difference between a system that can receive fixes and one that cannot.
2
Count what lives unmanaged in the production default solution
Anything customized directly in production sits in the default solution and cannot be moved as a unit, which means it is invisible to every deployment and will be silently absent from any environment rebuilt from source. The count of those components is the most honest single measure of whether this project has a release process at all.
3
List every writer against each column that matters
For the tables carrying the important logic, list everything that fires on create and on update: registered plug-in steps with their stages and execution order, real time and automated flows, classic workflows, and business rules. Where two or more of them write the same column you have found a race, and races are what produce the bug that only happens sometimes and can never be reproduced on demand. In a troubled project three writers on one column is ordinary rather than exotic.
4
Rank the last thirty days of failures rather than reading code first
Pull the plug-in trace log and the flow run history for the last month and sort by failure count. Nobody has been reading these, which is exactly why they are useful: they are an unedited record of what actually breaks, ranked by how often, and they usually contradict what the team believes about where the problems are. Start remediation from the top of that list rather than from the component that is most annoying to look at.
5
Separate synchronous work from work that should be in the background
Check which plug-in steps are registered synchronously and what they do. Calls to external systems inside a synchronous step are the usual answer to why the form takes twenty seconds to save and why saves fail whenever the other system is slow. Moving that work to an asynchronous step or a flow is often a day of work that changes the daily experience of every user, which makes it one of the highest value fixes available early in a rescue.
6
Find the code doing what configuration already does
Read the form scripts for field visibility toggling, requirement level changes, and simple arithmetic, all of which business rules, calculated columns, and rollup columns handle natively. This is usually the largest single block of removable code and the safest to remove, because the platform equivalent is declarative, visible to the next administrator, and survives updates without maintenance.
7
Check dependencies before deleting anything
A component that looks orphaned is often referenced by a view, a flow, a plug-in step, or a report that nobody thought to check. Run the dependency check on every candidate for deletion before it goes on the remove list. Skipping this is how a debt reduction exercise becomes the incident that proves the rescue made things worse.
How we fix it, phase by phase
Weeks 1 to 2
Make change safe before making any change
Nothing is refactored until there is a way to release and roll back. Separate development, test, and production environments, work done in unmanaged solutions in development and shipped as managed solutions, and the unpacked solution committed to source control so the state of the system exists somewhere other than the environment itself. Doing this first feels like a detour when everyone wants fixes, and it is the reason the fixes hold.
Weeks 2 to 3
Remove the pure risk before touching the load bearing debt
The debt inventory already separates the customizations the business genuinely runs on from the ones that are pure risk. Take the second group first: unreferenced scripts, disabled workflows nobody re-enabled, half finished components, and code the platform now does natively, each removed after its dependency check and each released on its own so a problem points at a single change. This is the fastest measurable reduction in the surface area available, and it takes no product decisions from anybody.
Weeks 3 to 5
Give every rule exactly one home
Where several components write the same column, decide which one owns it and retire the others deliberately rather than switching them off and leaving them registered. The ordinary shape of the answer is that validation belongs in business rules or a synchronous plug-in, cross record and external work belongs in an asynchronous step or a flow, and classic workflows are retired into whichever of the two fits rather than left running alongside their replacement. Execution order is set explicitly for what remains, and the intermittent bug goes away because the race that caused it no longer exists.
Weeks 5 onward
Rewrite the load bearing debt one component at a time
What is left is the code the business depends on and cannot lose. It is replaced behind the behaviour it already has, one component per release, with the old path removable and the change reversible. A big bang refactor of load bearing customizations is how a rescue becomes the next failed project, so this phase deliberately runs slowest and continues past the reset into ordinary support work, with the remaining inventory as the backlog.
What done looks like
Every important column has one writer, no component the business depends on sits in an unmanaged layer over a managed one, releases go from development to production through a repeatable path, and the failure counts in the trace log and flow history are falling week on week.
Integrations that fail quietly and lose records nobody notices
What it looks like
Nothing looks broken. No error appears in front of a user, there is no alert and no ticket, and yet finance keeps finding orders that never arrived and the support team keeps meeting customers whose details were updated in the other system a week ago. Quiet integration failure is the pattern that survives longest in a troubled project, because it produces no symptom until somebody reconciles, and on a project already in crisis nobody is reconciling.
How we confirm it, step by step
1
Find out whether a failure reaches a human at all
For each integration, trace what happens on a failed run. A flow with no failure path configured, a plug-in that swallows its exception, or a queue with no dead letter handling all fail invisibly by design. Ask one question of each interface: if this failed at three in the morning, who finds out and how. Where the answer is that somebody would eventually notice the data was wrong, the integration has no error handling regardless of what the design document says.
2
Count the failures already sitting in the run history
Read the run history and error logs for the last month for each interface and count what has already failed. On a stalled project this number is usually well into the hundreds and the records behind it have simply never arrived. That count is the size of the backlog to be replayed, and it needs to exist before anyone promises the business that the integration has been fixed.
3
Check who owns the connection and the application user
Connections owned by a named person break when that person leaves, changes password, or loses a license, and on a rescue that person has often already gone. Establish what each interface authenticates as, move anything running on a personal connection onto a service principal and application user with its own security role, and check that role is scoped to what the interface needs rather than set to system administrator because it was quicker.
4
Look for throttling rather than logic errors
Interfaces that work in testing and fail in production are frequently hitting service protection limits under real volume and returning throttling responses that the calling code treats as failures. Check the error detail for those responses before rewriting any logic, because the fix is retry with backoff and batching rather than a redesign, and rewriting logic that was never wrong is an expensive way to keep the same symptom.
5
Test whether replaying a message duplicates it
Take one message and send it twice into a sandbox. If the second one creates a second record, the interface is not safe to replay, which means the failure backlog cannot be recovered without generating duplicates and every retry so far has been making the data worse. Alternate keys and upsert behaviour are what make replay safe, and putting them in place comes before any recovery run.
How we fix it, phase by phase
Week 1
Make failure loud
Before anything is fixed, every interface gets a failure path that reaches a person: a configured error branch, a notification to a monitored channel, and a record of the failed payload that can be inspected and replayed later. Making failure visible first is what turns an unknown quantity of data loss into a list, and it is also what proves afterwards that the remediation worked.
Weeks 1 to 2
Make replay safe before replaying anything
Alternate keys on the tables the interfaces write to, upsert semantics instead of create, and retry with backoff on throttling responses. Only once a message can be sent twice without creating a second record does recovery of the backlog become possible, and this ordering is not optional: replaying into an interface that is not idempotent turns a data loss problem into a duplication problem on top of it.
Weeks 2 to 3
Recover the backlog and reconcile both directions
The failed payloads collected above are replayed in date order into a sandbox first, then production, and the result is reconciled against the other system in both directions rather than just checking that the records arrived. Reconciling one way finds what is missing here. Reconciling the other way finds what was sent from here and never landed there, which is the half most recovery exercises quietly skip.
Week 4 onward
Leave a reconciliation the business runs without us
The interface ends with ownership: a service principal rather than a person, a scheduled reconciliation report comparing counts on both sides, an alert when the difference exceeds an agreed tolerance, and written documentation of what each interface moves, in which direction, and who to contact on the other side. An integration nobody is measuring is one that will fail quietly again, and the rescue is only finished when somebody other than us would notice.
What done looks like
Every interface has a failure path that reaches a named person, replay is safe, the historical backlog has been recovered and reconciled in both directions, and a scheduled reconciliation makes the next divergence visible within a day rather than a quarter.
Naming which of these patterns you are actually in is the whole job of the audit, and getting it wrong is the most expensive mistake available in a rescue: a team convinced the problem is technical hires developers to build on a data model that was wrong before anyone wrote a line of code. Why the root cause has to be named before anything is fixed is set out in name the root cause before you fix anything, and the assessment method that separates the patterns, together with a worked remediation of Customer Service unified routing, is in Dynamics 365 rescue services: how to fix a failed or stalled implementation. For what reading and stabilizing a half built environment feels like from the inside, read taking over a Dynamics 365 project that went wrong.
Two of these patterns have their own deep dives on the site rather than being repeated here. For the mechanics behind the migration corrections above, including preserving created dates, the dependency order, files as their own workstream, and the reconciliation pack, read the practical data migration guide and, for high volume reloads, fast Dataverse bulk import. For the customization debt patterns themselves, and the audit passes that find unsupported form code, read common Dynamics 365 customization mistakes and how to fix them. Where you want this diagnosis run without a crisis attached to it, the same inspection is sold as a fixed scope Dynamics 365 health check and technical audit.
Process and Milestones: Our Fixed Scope Reset Phase
Most rescue offers describe the same undated diagram: assess, then stabilize, then optimize. The shape is right and it tells you nothing, because no phase has a date on it, no phase has an artifact attached to it, and a team that has already been told twice that things are on track has no way to check. Our first six weeks are a fixed scope reset phase instead: a fixed price, a fixed end date, and a named deliverable at every milestone, so immediate, visible progress is something you can point at rather than something we assert.
The reset phase has two inputs that everything else is built on, the initial triage and the stakeholder realignment, and then a milestone timeline you can hold us to.
Initial triage: the code and configuration audit
In the first week we read the environment rather than the story around it. We map the solution layers to see what is managed, what is unmanaged, and what was edited straight into the default solution in production. We find where business logic actually lives, which in a troubled project is usually three places at once: a plug-in, a real time flow, and a classic workflow all firing on the same record with nobody controlling the order. We read the Dataverse model and relationships, forms and business rules, JavaScript and custom APIs, PCF and canvas components, security roles, integrations, and the accumulated failures in flow and plug-in run history. In Field Service we also read requirements, resources, territories, working hours, locations, and the optimization configuration. In parallel we fix, in a sandbox, the failures that are hurting users today, so the first week produces relief and not just findings.
The technical debt inventory
Every finding from the audit becomes a line in a written inventory, and every line carries four things: what it is, what it blocks or breaks today, the risk of leaving it alone through the next release wave, and the rough effort to fix it. Each line is then marked keep, rework, rebuild, or delete, and separated into the customizations that are load bearing, meaning the business genuinely runs on them, and the ones that are pure risk, added for a reason nobody remembers and waiting to break. This inventory is what turns an opinion about the project into a decision the sponsor can make, and it belongs to you: it is written to be actionable by any partner or by your own team, including if you decide not to use us.
Stakeholder realignment
A stalled project has usually lost the agreement it started with. Sponsors want the original scope, process owners want the promises they were made, users want the spreadsheet back, and the remaining original team is defending decisions nobody has revisited in a year. We run short separate sessions with each group, write down the decisions the system actually has to support, and then bring the groups back together to agree three things in writing: what is in and out of the reset, what working means for the first module in language a user would recognize, and who the single decision maker is. We also fix a weekly checkpoint of thirty minutes with the same people and the same agenda, which is what stops the reset drifting back into the pattern that caused the crisis.
The reset timeline and what lands at each milestone
The windows below are the ones we commit to on a normal Dynamics 365 Customer Engagement or Power Platform takeover. Where an environment is unusually large or access takes longer to arrange, we move the dates before the reset starts rather than letting them slip afterwards.
Days 1 to 5
Access, triage, and the immediate risk list
Day one is access: security roles and environment permissions rather than a server handover, which is the one thing a cloud takeover makes easy. We then triage what is actively hurting people, the flow overwriting records, the form that will not save, the integration dropping rows, the schedule sending crews to the wrong side of the city. Those are reproduced in a sandbox copied from production, fixed, and promoted through a controlled path inside the first week. You get a triage log of what was found, an immediate risk list of what must not be touched until the audit is done, and visible relief for users within days rather than at the end of a discovery phase.
Deliverable: Triage log, immediate risk list, and the first emergency fixes released
Week 1 to 2
Environment Audit Report and Technical Debt Inventory
By the end of week two you have the two documents the rest of the rescue depends on. The Environment Audit Report states what is sound, what is broken, what is missing, and which root causes are producing the symptoms you can see, with the evidence from the environment behind each finding. The Technical Debt Inventory lists every item marked keep, rework, rebuild, or delete, split into load bearing and pure risk, with effort against each. If a design genuinely cannot be salvaged, this is where we say so in writing rather than layering more work on a broken foundation. Both documents are yours whatever you decide next.
Deliverable: Written Environment Audit Report plus the line by line Technical Debt Inventory
Week 2
Stakeholder realignment and a fixed scope for the reset
The audit and the inventory are walked through with the sponsor, the process owners, and the users in the realignment sessions, and the output is a single page everyone signs: what the reset will deliver, what it will deliberately not touch, the acceptance criteria for the first module in the words a user would use, the fixed price, and the date the reset ends. Fixed scope means the number and the end date are set before any of the build work starts, so the reset cannot become the same open ended overrun you are already living through.
Deliverable: One page Reset Scope with a fixed price, a fixed end date, and named acceptance criteria
Week 3 to 4
First working module delivered
The first increment is one module that a team uses end to end, chosen with the sponsor for the shortest distance to visible value: a case to resolution path with routing that finally works, a sales process that matches how deals are actually run, or a week of Field Service bookings the dispatch team is willing to trust. It is built in a sandbox against real data and real volume, tested against the process it supports rather than a script, and promoted as a managed solution. By the end of week four the people who lost confidence in this project have something in front of them that works, which is the moment a rescue stops being a promise.
Deliverable: One end to end module live in your environment, tested against the real process, with real users on it
Week 5 to 6
Second increment, a clean release path, and the exit decision
The second increment proves the first was not a one off, and alongside it the environment moves onto proper solution based application lifecycle management: separate development, test, and production environments, changes made in unmanaged solutions and promoted as managed ones, so a change can be reviewed and rolled back instead of typed live into the system people depend on. You also get the handover pack: how the solution is built, why it is built that way, what was fixed, and what remains on the debt inventory. The reset then ends on its agreed date with a decision that is yours: continue on a longer plan with the remaining inventory as the backlog, hand the work back to your own team, or stop with a documented, stable environment and a report any partner can act on.
Deliverable: Second module delivered, development, test, and production solution path in place, written handover pack, and a go or stop decision
A stalled migration is its own kind of failure and it needs its own answer. The project has not cut over, both systems are still live, people are entering the same record twice, and the date keeps moving without anyone being able to name the thing that is actually blocking it. That is not the same problem as a migration that finished and brought the wrong data across, which is set out as its own pattern in a data migration that brought the wrong data across. The two have different deadlines, and treating one as the other is how a rescue wastes its first month.
The deadline on a stalled migration does not belong to us and it does not belong to your project plan. It belongs to the system you are moving off: the date it is switched off, the renewal it does not survive, or the day read access ends with a contract that belongs to somebody else. Everything below is ordered around finding that date first. We take this work on for Dynamics 365 Customer Engagement and Dataverse migrations, from Salesforce, from an on-premises Dynamics CRM estate, or from a line of business system somebody built in house, either directly or on a white-label basis for other Microsoft partners.
What a stalled migration actually looks like
These are the states we find a halted migration in. Most projects are in three or four of them at once, and the last one is the only one with a clock on it.
The cutover date has moved twice and nothing about the reason changed
A first slipped cutover is ordinary. A second and a third date set while the blocking item is still the same unreconciled load, the same unsigned mapping, or the same set of lookups that never resolved means nobody has diagnosed why the migration stopped, so the new date is a hope rather than a plan. This matters more on a migration than on a build, because the system you are moving off usually has a contract renewal or a decommissioning date that does not move when the project does.
Both systems are live and people are entering data twice
Dual running was meant to last a fortnight and has lasted a quarter. One team works in the old system because that is where the history is, another works in Dataverse because that is where the new process lives, and nobody can produce one report across both. Every extra day of dual running writes more records into the source that the migration has to pick up again, so the load that was tested is no longer the load that has to run.
The load has run but no reconciliation was ever produced
Counts were reported verbally in a status meeting and never written down table by table, or a single total row count was compared, passed, and treated as proof while ownership, state, currency, created dates, and lookups were never checked at all. A migration with no reconciliation pack is not finished, it is unverified, and until somebody produces the per table numbers nobody can say whether cutover is a week away or a rebuild away.
The mapping exists only in the head of somebody who has left
The field and option set mapping lives in a spreadsheet with three versions and no owner, or only inside the migration tool itself, and the person who made the judgement calls about which values collapse into which, what counts as a duplicate, and who owns the records with no owner is no longer on the project. The load can still be run. Nobody can say what it means, which is why it is not being run.
The migration is a set of scripts that only runs on one laptop
The extract and the load were built as one time scripts against somebody's personal credentials rather than as something rerunnable. It worked once, in an order nobody wrote down, and it cannot be run again without the person who wrote it. A load that cannot be rerun cannot be rehearsed, and a migration that cannot be rehearsed will only ever be tested in production, on the day, in front of the business.
The read window on the source system is closing and nobody has said so
The old system is being switched off, a licence lapses at the next renewal, or read access ends with the outgoing vendor's contract. Anything that has to be extracted again has to be extracted before that date, and after it the difference between missing and permanently lost stops being recoverable. This is the single item that sets the deadline for every other decision in the rescue, and it is the one most often absent from the status report.
The migration takeover, week by week
Six weeks from the day access is granted to a rehearsed cutover on a typical Customer Engagement estate. The first two weeks are fixed scope and end in a Migration State Report you keep whatever you decide next, so you are not committing to the whole thing to find out where you stand.
Days 1 to 3
Stop the writing, secure both ends, and find the real deadline
Anything still loading on a schedule is turned off before anything else is discussed, because a migration that is still writing will undo corrections as fast as they are made and it is the most common reason a data cleanup runs twice. We then preserve evidence at both ends: a copy of the target environment, the staging tables, the migration job logs, and confirmation of how long the environment keeps restore points. Last, we establish the date the source system stops being readable from the contract or the decommissioning plan rather than from the project plan, because that date, and not our schedule, sets the deadline for anything that has to be extracted again.
Deliverable: Delta loads paused, a copy of production and the staging data preserved, and the date the source stops being readable, in writing
Days 3 to 8
Reconcile what has already landed
This is the count nobody has produced. For every table already loaded we compare the target against the source and check the five things a total row count hides: created dates clustered on the load date instead of spread across the years the source covers, records owned by the migration account instead of by people, closed records that arrived open and are inflating the backlog, money columns that fell back to the base currency, and lookups that failed to match and were written as empty or dropped into a text column. This step turns opinions about the migration into a number, and the number usually changes what everyone believed the remaining work was.
Deliverable: A per table reconciliation of target against source covering row counts, created dates, ownership, state and status, currency, and unresolved lookups
Week 2
Recover the mapping and get the open decisions signed
We rebuild the mapping from what is actually in the source and the target rather than from whichever spreadsheet version turns up, then take the unmade judgement calls back to the people who can answer them: which option set values collapse into which, what happens to records whose owner has left, what counts as a duplicate, and how much history is genuinely in scope. These are business decisions a migration team cannot make on your behalf, and leaving them unmade is the most common reason a migration stalls with no technical blocker anyone can name. The Migration State Report is yours whatever you decide next.
Deliverable: One current field and option set mapping with the deduplication and ownership rules written into it, signed by the business, plus a Migration State Report
Week 2 to 3
Rebuild the load as something that can be run again
The scripts become a load that somebody on your side can run: a fixed dependency order so parents exist before children, a service identity instead of a named person's account, the created on override set on every table where ageing or service level reporting depends on it, alternate keys in place on every table an interface writes to, and errors written out per table rather than scrolling past in a console. This is not polish. Rehearsal is what removes almost all cutover risk, and nothing that cannot be rerun can be rehearsed.
Deliverable: A rerunnable load running from a service identity on infrastructure you own, in a documented dependency order, with per table error output
Week 3 to 5
Rehearse the load twice against full volume
The load runs into a sandbox at full volume, is reconciled, corrected, and then run again, because two loads is the normal number rather than a sign of caution. Files, notes, and activity history are treated as their own workstream and counted in bytes as well as rows, since a partial file migration reports a healthy row count while the documents themselves are truncated or absent. The second run is timed end to end, so the cutover window is planned from a measurement rather than an estimate, and the runbook falls out of it: the order, the duration of each step, who does what, and the point after which the decision is to go forward rather than back.
Deliverable: Two rehearsal loads into a sandbox, each closed by the same reconciliation, and a timed cutover runbook
Week 5 to 6
Cutover with a rollback, then hypercare on the data
Cutover runs the rehearsed sequence with a defined rollback point, and the source system goes read only rather than being switched off. The business signs the same reconciliation it has now seen twice, so acceptance is a comparison rather than an assertion. Dual running ends on an agreed date instead of by drift. And we hand over the honest third list: the data that is in neither system and has to be either accepted or recaptured by hand. Writing that list down and telling the business what is on it is what stops a migration being argued about six months later.
Deliverable: Cutover executed against the runbook, the reconciliation pack signed by the business, dual entry ended on a date, and a written list of what was not brought across
The three verdicts a stalled migration ends in
The reconciliation in the first week is what decides which of these you are in, which is why it happens before anyone quotes for the rest. We would rather tell you in week one that the migration needs re-scoping than bill you for four more months of finishing something that was never going to land.
The verdict
When this is the answer
What it costs and what it saves
Resume the existing migration
The reconciliation comes back clean on the tables already loaded, the mapping can be recovered and signed, and the load can be made rerunnable without being rewritten. This is more common than a team in the middle of a stalled migration expects, because the real blocker is usually an unmade business decision rather than broken engineering.
The cheapest outcome, measured in weeks, and everything already paid for stays paid for. What it costs is honesty about the decisions that were skipped, because resuming without making them stalls the migration a second time at exactly the same place.
Rebuild the load, keep the mapping
The mapping is sound but the load cannot be trusted or rerun: created dates were never overridden, ownership sits on the migration account, or the extract exists only as one time scripts. The data in the target is wrong in ways that can only be corrected by loading it again.
The middle outcome and the most frequent one. The analysis survives and only the engineering is redone, which is a matter of weeks rather than months. It depends on the source still being readable, which is why that date is established on day one rather than discovered in week five.
Re-scope the migration
The reconciliation shows damage across most tables, users have already worked on the migrated records so a straight reload would destroy real work that exists nowhere else, or the mapping was built against a Dataverse model that does not fit the business. The question stops being how to finish this migration and becomes how much history the business actually needs.
The outcome nobody wants to hear and the one that most often saves the project. Cutting scope to open records plus the accounts and contacts behind them, with the rest kept readable in the source or in an archive, removes the second load, the second reconciliation, and most of the remaining timeline.
The mechanics behind these steps are written out in full elsewhere on the site rather than repeated here. For preserving created dates, the dependency order, files as their own workstream, and what a reconciliation pack contains, read the practical data migration guide, and for the design decisions behind a Salesforce move, Salesforce to Dynamics 365 migration best practices. Where the rehearsal loads are large enough that throughput is the constraint, fast Dataverse bulk import covers how we run them. And where you want the reconciliation run on its own, without a takeover attached to it, the same inspection is sold as a fixed scope Dynamics 365 health check and technical audit.
Rescue & Takeover in Armenia: Our Local Advantage
If your Dynamics 365 project has failed or stalled in Armenia, the team that recovers it is already here. Solzet is a Dynamics 365 Customer Engagement and Power Platform consultancy in Yerevan, so an Armenian client gets an engineer on site the same morning, a rescue run on the same working day and the same holiday calendar as their own company, and a partner who already knows the systems on the other side of the integrations. The 72 hour emergency takeover process further down this section is the local path through it, hour by hour.
Most recovery advice written for this problem is written for nobody in particular: a generic assess, stabilize, optimize sequence that would read the same in any country. What follows is the opposite. It is only the part that changes because the client and the rescue team are in the same city. The platform method itself does not change and is not repeated here: the emergency sequence is in the first 72 hours flowchart above, the assessment discipline behind it is set out in Dynamics 365 rescue services: how to fix a failed or stalled implementation, and the same inspection run before anything is on fire is our Dynamics 365 health check and technical audit, which is the right starting point if your project is drifting rather than burning.
Not an overlapping working day, the same working day
A distant partner sells you hours of overlap. Inside Armenia there is nothing to overlap, because both teams start at nine and finish at six on the same clock, break for the same public holidays, and lose the same week at the start of January. On a healthy project that is a convenience. On a rescue it is the whole thing, because the largest risk on a failing implementation is decision latency rather than engineering: the question asked on Tuesday that gets answered at next month's steering meeting. A question asked here at eleven is answered before lunch, and a freeze that needs the general director's agreement gets it in an afternoon rather than in a thread.
An engineer at your desk the same morning
Lockdown day goes faster with a person in the room, because the things that block it are people problems: an administrator password nobody has written down, a laptop that holds the only working connection to a local system, a colleague who will show you the real process but will not write it in a ticket. We are in Yerevan, so an engineer can be at any office in the city inside the hour, and Gyumri, Vanadzor, and most regional centres are a half day by road rather than a flight, a visa, and a travel budget. On site is a normal working decision here rather than an event that has to be justified.
How decisions and sign-off actually work here
In most Armenian companies authority sits with the founder or the general director rather than with a steering committee, agreements are made verbally in a meeting and treated as final, and the paperwork follows later as acceptance acts your accountant needs for the books. A partner who does not know that either waits for a governance structure that is never going to exist, or takes a verbal yes and has nothing to point at three weeks later. We work the way the company works: sit in the room, take the decision there, and put it in writing the same evening so the record exists without anyone being asked to change how they run their business. Being an Armenian company ourselves also means the contract, the acts, and the invoicing sit under the same law your finance team already works under, with no foreign counterparty on the other end of an emergency.
The systems on the other side of your integrations
A failed implementation in Armenia is rarely broken only inside Dynamics 365. It is broken where it touches 1C or an Armenian Software accounting system, a bank client, ArCa card processing, or the State Revenue Committee portal the company files through, and those interfaces are the ones where a bad write turns into a filing problem rather than a data problem. We read that estate as ordinary requirements because we work in it every week, and their owners are usually one accountant or one contractor in this city who can be reached by phone the same afternoon rather than a vendor support queue in another country.
Where your tenant and environments actually sit
There is no Microsoft datacenter in Armenia, so an Armenian tenant's environments live in a foreign region that was chosen, usually without thought, when somebody signed up, and that choice follows the tenant afterwards. It is worth knowing which one you are on before an emergency rather than during it, because it sets your latency, it sets the conversation about where personal data is processed under the Armenian data protection law, and it adds a second obligation if you serve customers in the EU. Armenia's international connectivity also runs overland, which shows up as a slow client rather than as an outage and is regularly misdiagnosed as a slow build.
A small market, where the team that built it is still reachable
The Armenian Microsoft market is small enough that the developer who built your system is usually still in Yerevan, often at a company we already know, and frequently still willing to answer the phone. That hour is the cheapest thing in the entire rescue and it recovers what no audit can: which customization was a deliberate decision and which was a workaround. The same smallness cuts the other way, because everyone will meet again and nobody wants to be blamed in public, so we keep the audit about the environment and never about the people, and we ask for facts in writing rather than an account of what went wrong.
The 72 Hour Emergency Takeover Process for Armenia
Read it top to bottom. Each box is either something we do or a question whose answer changes what happens next, and the two branches under a question are what happens in either case. The windows are Yerevan time, which for a client in Armenia is simply the time, and the clock starts when access is granted rather than when a contract is signed. Everything here is the local path: what an engineer in the room changes, the tenant question that has to be settled before anything commercial is discussed, which integrations get frozen first, and how a verbal decision becomes a written record the same evening.
1Hour 0, Yerevan timeStep
The call, answered by the person who will run it
You reach an engineer in Yerevan inside your own working day, not an account manager in another region who will schedule a call for next week. Two people are fixed on that call: someone on your side who can authorize access to the environment, which in an Armenian company is usually the founder or the general director and is therefore reachable directly, and a named person on ours who owns the intervention until the stabilization statement is written. The clock below starts when access is granted rather than when a contract is signed, because the contract can be finished while the environment is already being made safe.
?Hour 0 to 2Decision point
Are you in Yerevan or in the regions?
This decides whether the first day is run from your office or from ours, and it is asked early because being in the room changes what gets done on day one, not just how it feels.
In Yerevan
An engineer is at your office the same morning while the remote lockdown runs in parallel. That is the fastest way through the things that actually block a first day: the administrator credentials nobody wrote down, the one laptop that holds a working connection to the accounting system, and the colleague who will demonstrate the real process but will not describe it in writing.
Outside Yerevan
Lockdown starts remotely inside the hour and travel is scheduled where it earns its place, usually for the audit rather than for hour one. Gyumri, Vanadzor, and most regional centres are a half day by road, so on site attendance is a decision about value rather than about budget, and it can be repeated later in the rescue instead of being spent once.
3Hour 0 to 8Step
Environment lockdown, with someone in the room
The platform mechanics are the same everywhere and are set out in the general flowchart above: production frozen, administrative roles cut back, direct edits stopped, and a sandbox copy taken so the failing state is preserved. What being local adds is speed through the human part of it. Roles are cut back with the people who hold them sitting there rather than through a chain of forwarded emails, and the person whose access is about to be reduced hears why from an engineer in the same room, which is what stops a lockdown turning into a fight on the first afternoon.
?Hour 2 to 6Decision point
Whose tenant is your data actually in?
This question matters more in a small market than anywhere else, and it is asked before any commercial conversation with the outgoing party, because the answer decides whether this is a takeover at all.
Your own tenant
This is an ordinary vendor takeover, and it runs on the order and the clock set out in the vendor takeover list above. The tenant is yours, so administration can be recovered whatever the state of the relationship, and nobody on the other side has to agree to it.
A contractor's or a reseller's tenant
It happens here, usually because a small local developer built the first version inside their own subscription and it was never moved. That is not a takeover, it is a migration, and we say so on day one rather than in week three. The first action is securing a full export of your data and the unpacked solution while access still exists, because the only leverage in the conversation that follows is the copy you already hold.
5Hour 4 to 24Step
Freeze the local integration surface before anything cosmetic
Triage here starts with the interfaces where an error becomes a filing problem rather than a data problem: the link to 1C or an Armenian Software accounting system, the bank client, ArCa card processing, and whatever posts to the State Revenue Committee portal. Anything writing into those is stopped first and reconciled before it is turned back on. Their owners are usually one accountant or one contractor in this city, so the conversation that a foreign partner books as a vendor ticket happens here as a phone call the same afternoon, and the reconciliation is agreed with the person who will sign the filing.
6Running from hour 0Runs in parallel
One written brief a day, in the language your sponsor reads
Decisions in an Armenian company are frequently made verbally and treated as settled, and that is a workable way to run a business right up until a project fails and nobody can say what was agreed. We do not fight it. The decision is taken in the room, and the same evening it goes out in writing in Armenian, Russian, or English, whichever the decision maker actually reads: what was found, what changed, what happens next. Users get a plain note about what is frozen and what to do meanwhile, because in a company of this size silence at six in the evening is what starts the rumour by nine the next morning that the system is being switched off.
7Day 2 to 3Step
One hour with the people who built it, in person
The market is small enough that the previous developer is usually still in Yerevan and often still answering the phone, so the handover hour that is theoretical in most rescues is genuinely available here. We ask for one hour with a prepared list rather than an open conversation, and we ask it in the language the specification was written in, which locally is Russian at least as often as English. We keep it about the system and never about fault, partly because that is what gets the answers, and partly because in a market this size everyone involved will work together again.
8Hour 72, end of day threeOutcome
The stabilization statement, and the decision that is yours
One page, in the language your board reads: what was locked down, what was fixed, what is still dangerous, which local interfaces remain frozen and what that means for the next filing date, what data damage is recoverable and what is not, and what it would take to move from stable to working. It is yours whichever way you decide, and it is written to be usable by your own administrator or by another Yerevan firm, not only by us. If you continue with us, it becomes the front door to the fixed scope reset phase set out above, with a fixed price, a fixed end date, and a named deliverable at every milestone.
After hour 72 an Armenian rescue rejoins the ordinary path on this page: the vendor takeover of the keys, the Environment Audit Report and Technical Debt Inventory at week two, and the fixed scope reset phase that puts a working module in front of real users by week four. How the firm behind that is structured, staffed, and certified is set out in our company profile. If you are outside Armenia but inside the same working day, the wider regional case, including the handover mechanics and how a Yerevan team compares with a distant global consultancy, is the section immediately below.
The Takeover Process for CIS & Caucasus Businesses
If you are looking for a Dynamics 365 takeover specialist in a similar time zone to your own team, the practical answer for the CIS and the Caucasus is a partner in the region rather than one flying in from Western Europe. Solzet works from Yerevan in GMT+4, and Armenia does not observe daylight saving, so the offset never moves through the year: the same clock as Tbilisi and Baku, one hour ahead of Minsk and Moscow, one hour behind Tashkent and Almaty, two behind Bishkek. From Kaliningrad to Kyrgyzstan that is a full working day of live overlap rather than a few hours at the edges, which is the difference between a rescue that makes decisions daily and one that makes them overnight.
The steps below are the handover mechanics of a takeover: how control of the environment actually moves, how knowledge is extracted from the team that is leaving, and where the audit and stabilization work described above meets the reality of a project built in this region. Solzet is a Dynamics 365 Customer Engagement and Power Platform specialist rather than a full range integrator, and our company profile sets out how the firm is structured and staffed.
1
Environment and tenant access, day one
The first blocker in a takeover is almost never technical. It is who holds the keys. On a stalled project we routinely find the outgoing partner holds Global Administrator on the Microsoft 365 tenant, owns the CSP licensing relationship, is the named owner of the Power Platform environments, and controls the Azure app registrations the integrations authenticate through. Day one is therefore a written access list rather than a workshop: Power Platform Administrator and Dataverse System Administrator for the people who will now own the system, ownership of the service principals and application users the integrations run under, the Azure DevOps or GitHub repositories where the solution source actually lives, and the deployment service connections. Because a Dynamics 365 environment is cloud tenant property, you can insist on this even where the relationship has broken down, and it is the part clients most often do not realize is theirs to take. We run that conversation live with your IT people in your own working hours, which for a team anywhere between Minsk and Tashkent means it is finished the same day rather than after two overnight email cycles.
2
Knowledge transfer from the outgoing team, in the language it was built in
The handover window with a departing partner or a leaving internal developer is usually short and it is the cheapest hour of the entire rescue, because it answers the questions no audit can recover: which customization was a deliberate design decision and which was a workaround for something else, what was promised verbally to which department, which integration nobody dares touch and why. Across the CIS and the Caucasus that conversation happens in Russian at least as often as in English, the technical specification and the acceptance protocols are written in Russian, and the comments inside the plug-in code frequently are too. We work in Russian, Armenian, and English, so none of it passes through a translation layer and nobody on your side is asked to interpret their own project back to their new partner. Where the outgoing team has already gone, this step becomes a reconstruction from the environment itself, which is slower and is exactly why the handover hour is worth fighting for.
3
The code and configuration audit
From here the work is the same discipline we apply on any takeover, and it is set out in full above rather than repeated here: we read the solution layers, where business logic actually lives, the Dataverse model, security roles, integrations, custom code, and the accumulated failures in flow and plug-in run history, and every finding becomes a line in the Technical Debt Inventory marked keep, rework, rebuild, or delete. What is regional about it is the estate around the edges. The system on the other side of an integration in this region is often 1C, a local bank client, or a national tax or e-invoicing service rather than SAP, and a build that has been running for two years has usually accumulated local process assumptions that were never written down. We treat those as ordinary requirements to be read, not as exotic scope to be discovered later.
4
The stabilization plan and the reset
The stabilization methodology above applies unchanged: lock the environment down, stop the damage, diagnose under pressure, put one honest line of communication in place, and then move into the fixed scope reset phase with a fixed price, a fixed end date, and a named deliverable at every milestone. Distance is what changes the value of that, not the method. Emergency stabilization only works if a decision can be made on the day it is needed, and a partner four or five hours behind your working day cannot give you that. Yerevan sits inside your day, so the daily written brief lands in your morning and the person who wrote it is available to be argued with in the afternoon.
5
Handover back to your own team
A rescue is finished when your people can run the system without us, which means the documentation has to be readable by the administrators who will actually use it. We write the handover pack, the solution and deployment approach, and the configuration notes in the working language of the team that inherits them, and we train those people directly rather than through an account manager. If you keep us on afterwards it is the same engineers on support, not a separate desk in another region.
Why a regional partner is the stronger choice for a takeover
This is not an argument that global consultancies are bad at Dynamics 365. Firms such as Avanade run strong practices, and for a multi country programme with heavy governance obligations they are frequently the right answer. A takeover is a different job, and the things it depends on are the things distance takes away.
On a takeover
Regional partner in Yerevan (Solzet)
Distant global consultancy
Overlap with your working day
Yerevan is GMT+4 with no daylight saving, so the offset never moves: exactly the same clock as Tbilisi and Baku, one hour ahead of Minsk and Moscow, one behind Tashkent and Almaty, two behind Bishkek. A full working day of live overlap, every day of the year.
A Western European or US delivery team gives you a few hours at the edges at best, and an offshore delivery centre behind the sales office gives you an overnight cycle. A question asked at 15:00 is answered tomorrow.
The language the project runs in
Russian, Armenian, and English, from the engineers themselves. Specifications, acceptance protocols, handover sessions, and user training happen in the language your team already uses.
English only in practice, with the regional language handled by a local subcontractor or by your own staff acting as interpreters between their colleagues and the build team.
Getting someone into the room
Yerevan is roughly an hour from Tbilisi, under three from Moscow and Dubai, and Armenian consultants travel visa free across the EAEU and the wider CIS. A workshop or a go live can be attended in person at short notice.
Visas, long haul flights, and a travel schedule agreed weeks ahead. On site presence becomes an event rather than something that can be arranged for the week a rescue needs it.
What that travel costs you
Short regional flights and no billed long haul travel time. On site attendance is a line item you can afford more than once in a project.
Billed travel time, international flights, and per diems on top of a rate that is already several times higher, which is usually why the on site presence quietly stops after the kickoff.
Reading the estate around Dynamics 365
Integrations to 1C, local bank clients, and national tax and e-invoicing services are ordinary requirements here, and so are the local process habits a build accumulates.
A reference architecture built around SAP or Microsoft finance systems, with the regional reality scoped as an exception once it is discovered.
Who is actually on the build
Senior, Microsoft certified engineers, typically a team of two to three, and the people who diagnose the project are the people who fix it.
A staffing pyramid with senior architects on the pitch and a larger junior bench on delivery, plus an account layer between you and the engineers.
Where a global firm genuinely wins
A specialist team of this size is the wrong answer for a multi country programme spanning finance, supply chain, and CE with heavy governance obligations.
Exactly that: multi country scale, formal governance, and the ability to staff several workstreams at once across several Dynamics 365 product families.
The regional side of how we work is covered in more depth elsewhere on the site. For the consultancy models open to a buyer in Armenia and the Caucasus, and a vendor neutral due diligence checklist to apply to any of them including us, read the Yerevan local delivery guide. For how the partner relationship works day to day, the languages we deliver in, and how contracting, IP, and data protection are handled, see Microsoft Dynamics 365 partner in Armenia. And for the assessment method behind the audit step above, including the root cause families we test for, read Dynamics 365 rescue services: how to fix a failed or stalled implementation.
What we take over and rescue
We stay strictly within Dynamics 365 Customer Engagement and the Power Platform, the areas we build in every day, so a takeover is finishing work we know, not learning it on your budget.
Dynamics 365 Customer Engagement apps
Sales, Customer Service, and the wider Customer Engagement suite, including model driven apps, forms, business process flows, security roles, and the Dataverse data model underneath them.
Field Service and Resource Scheduling Optimization
Field Service work order, booking, and scheduling setups, including Resource Scheduling Optimization. We correct the inputs the engine actually depends on so it returns schedules dispatchers can trust: technician skills and characteristics, resource and requirement records, territories, and accurate locations.
Power Platform solutions
Power Apps canvas and model driven apps, Power Automate flows, Power Pages sites, and the integrations and connectors that tie them to Dynamics 365 and the rest of your systems.
Custom code and PCF controls
Plug-ins, custom APIs, JavaScript, and PowerApps Component Framework controls that were built badly, left unfinished, or are no longer maintainable, rewritten to a standard your team can support.
Data model and migration issues
A Dataverse schema that does not fit the business, duplicate or dirty data, and migrations that stalled or brought across the wrong things, straightened out so the platform reflects how you actually work.
Documentation and handover for an undocumented build
Solutions that work in some fashion but that nobody can explain, reverse engineered and documented so the current design, customizations, and dependencies are written down and the system stops being a black box.
Standard Implementation Timeline: 8 to 12 Weeks in Armenia
A Dynamics 365 Customer Service implementation takes eight to twelve weeks for a single support team, measured from a signed scope to a live desk with hypercare finished. Eight weeks is one team on one channel with no migration and no custom code. Twelve weeks is the same desk with a real data load, a live channel, and integrations to systems somebody else owns. Anything materially shorter is a pilot being described as an implementation, and anything materially longer usually has one of the eight causes listed further down this section already at work in it.
The plan below is what we actually deliver from Yerevan, not a phase diagram. Product documentation and most partner pages answer this question with an undated sequence of phases, which is the advice that produces the projects the rest of this page exists to rescue. What follows is the twelve week case written week by week, the six conditions that compress it to eight, the factors that cause overruns with the week each one usually becomes visible, and then the honest contrast: what the number means when the timeline has already been blown, which is the situation most readers of this page are actually in.
The region matters to this number more than it looks. Yerevan is GMT+4 with no daylight saving, so a business anywhere between Minsk and Tashkent gets a full working day of live overlap, and the largest schedule risk on a healthy implementation is decision latency rather than engineering. We work in Russian, Armenian, and English, so specifications and acceptance protocols do not wait on a translation layer, and Yerevan is roughly an hour from Tbilisi and under three from Moscow and Dubai, which makes being in the room for the pilot week affordable rather than an event planned a quarter ahead. The estate around Dynamics 365 here routinely includes 1C, a local bank client, and national tax and e-invoicing services, and we plan those integrations in week two instead of discovering them in week seven.
The twelve week plan, week by week
The phases overlap on purpose. Development starts while configuration is still finishing, and testing starts before the last load is done, because a plan where each phase waits for the previous one to close completely is a plan with a fortnight of idle time inside it. What does not overlap is the order within configuration: each step depends on the one before it being genuinely finished, and that order is set out step by step in the Customer Service and Omnichannel implementation guide below rather than repeated here.
Weeks 1 to 2
Discovery: the decisions, the environments, and what the data really looks like
Week 1
The eight planning decisions set out further down this page are answered and written on one page, and the development, test, and production environments are stood up with the solution and release path settled before anything is configured. A single decision maker is named and a thirty minute weekly checkpoint is fixed with the same people and the same agenda. This is also the week we ask for the thing most projects skip: half a day sitting beside two agents while they work real cases, because a queue topology drawn in a workshop and a queue topology drawn from watching the work are rarely the same document.
Week 2
The case model, its relationship to account and contact, and the governed option sets that will later carry your reporting. In parallel we profile the data you intend to bring across rather than accepting a description of it: row counts per table, the counts on the subset the business actually cares about, how the created dates are distributed, who owns the records, and how much duplication the source already contains. We also write down every integration, which system owns each one, and what it authenticates as. Both of those lists are cheap in week two and are the two most expensive things to discover in week seven.
What we need from you: Half a day each from the service manager and two working agents, read access to whatever holds the cases today, and a named person who can approve a decision inside the working day rather than at a monthly steering meeting.
Where this phase slips: If the list of queues and their owners cannot be written down because nobody agrees who owns what, that is a finding about the organization rather than about the platform, and it has to be resolved here. Encoded into routing instead, it comes back in week nine as a routing problem that no rule change fixes.
Weeks 3 to 5
Configuration: intake, then routing, then the clock
Week 3
Intake is proven before a single routing rule exists. The mailbox is approved and enabled for server side synchronization and confirmed as genuinely receiving, automatic record creation rules turn inbound mail into correctly shaped cases, and the queues decided in week one are created with their members and their assignment method. Proving intake first is not fastidiousness. A mailbox that quietly failed test and enable stops email routing before any rule is evaluated, and a team that has not separated the two loses most of a week debugging rules that were never reached.
Week 4
Unified routing, in the order the platform forces: the workstream with its type, distribution mode, and capacity model, then the work classification rule sets that stamp the attributes routing will read, then the route to queues decision list evaluated top to bottom with narrow rules above broad ones and a catch-all at the bottom pointing at a fallback queue somebody actually watches. Routing is then read through the diagnostics on real traffic rather than judged from the screen, because a rule that looks correct and never fires is almost always testing an attribute that was never stamped.
Week 5
The customer service calendar and business hours first, since every duration is computed against them, then the SLA KPIs for first response and resolution with the pause behaviour that was agreed in writing in week one, then warning and failure actions that reach a person. Entitlements are configured here or explicitly deferred with a date on the roadmap. Knowledge is seeded in the same week from the top case reasons in the last quarter of real traffic, with an owner and a review date on every article, rather than from a workshop wish list.
What we need from you: A tenant administrator available for the mailbox approval, sign off on the SLA pause conditions in the words a service manager would defend to a customer, and one subject matter expert per queue to check the classification attributes are the ones the desk really sorts on.
Where this phase slips: Configuring routing before intake is proven, or writing SLA KPIs before anyone has agreed what the clock measures, are the two ways this phase turns three weeks into five. Both feel like progress at the time, which is why they are so common.
Weeks 5 to 8
Development and data: the code that survived justification, and the loads that get reconciled twice
Weeks 5 to 6
The agent workspace and the live channel: session and workspace behaviour, what closes after wrap up, what survives a refresh, presence and capacity profiles, notification timeouts, and the customer inactivity warning and closure values, every one of them checked against how long a real conversation in your business takes rather than left at the default. Alongside it, the small amount of custom development that survived being justified one component at a time: a plug-in or custom API where the platform genuinely cannot do the thing, and a PCF control only where a supported form cannot. On a healthy build this list is short, and keeping it short is most of why the twelve weeks holds.
Weeks 6 to 7
The integrations are built against the inventory written in week two, each one with a failure path that reaches a person, alternate keys so a replay updates rather than duplicates, and retry with backoff where service protection limits apply. The first full data load runs into the test environment, in dependency order: users and teams, then ownership, then currencies and reference data, then lookups and parents, then the transaction rows. It closes with a reconciliation per table rather than a statement that the load worked.
Week 8
The second load, run against the corrections from the first, with the same reconciliation repeated so the numbers can be compared rather than asserted. Duplicate detection rules are published and scheduled and the alternate keys are in place on every table an interface writes to, which is the configuration that stops the same duplication returning through the same door a quarter after go live. Two loads is not caution, it is the normal number. A migration that is only run once is run for the first time in production.
What we need from you: Someone who can answer for each source system, and a business owner per table who will sign the reconciliation. The second is the one projects underestimate: a reconciliation nobody signs is a spreadsheet, not an acceptance.
Where this phase slips: This is the phase with a hard external deadline, because a source system is often switched off or made read only on a date somebody already agreed. Anything that has to be re-extracted has to be re-extracted before that date, and records loaded without the created date override cannot be corrected afterwards, only reloaded.
Weeks 8 to 10
Testing and pilot: the awkward cases, then real traffic with real authority
Week 8
Scripted testing proves the build does what it was configured to do, and then the awkward cases prove it does what the business needs: a case that crosses a weekend, one paused with the customer for two days, one reassigned to another queue halfway through its SLA, one arriving from a customer whose entitlement has run out. Those four are the cases a service manager will be asked about within a month of go live, and they are the ones a script never covers.
Week 9
A named pilot group works live traffic with the authority to change the build before it reaches anybody else. Routing diagnostics, SLA outcomes, and knowledge search results are read daily rather than at the end of the week. It is normal and expected for the queue topology decided in week one to change here, because this is the point at which it meets the people who work it. A pilot whose feedback cannot alter the configuration is a rehearsal, and it is why so many desks go live with problems everybody already knew about.
Week 10
The pilot findings are built, the regression pass runs over everything the changes touched, and performance is checked at real volume rather than at test volume, including the views and dashboards agents will keep open all day. The phase closes with a written pilot log of what changed as a result and a go or hold decision. Hold is a legitimate outcome here and is far cheaper than the alternative, which is discovering the same thing in week eleven with the whole desk on it.
What we need from you: A pilot team of real agents released from part of their normal workload, and a service manager willing to let the pilot change the build rather than defend the design.
Where this phase slips: Testing scheduled as the buffer at the end is how a twelve week plan becomes a sixteen week one, because it is the only phase whose scope is set entirely by what the earlier phases got wrong. Where testing is being compressed to protect a date, the date has already moved and nobody has said so yet.
Weeks 10 to 12
Go live and hypercare: the baseline first, then the cut over, then two weeks of somebody answering
Weeks 10 to 11
The current numbers are captured before go live, so that any claim about improvement later has something honest to compare against. The final delta load runs, the reconciliation is signed for the last time, and the desk cuts over with the previous route of last resort deliberately closed rather than left quietly available. A shared mailbox that stays open on cut over day is the single most reliable predictor of a desk that never fully adopts.
Weeks 11 to 12
Hypercare runs for two weeks minimum with a named person answering agent questions in the channel agents already use and a written daily note of what changed, so nobody has to ask twice. In the same fortnight the build is documented as it actually is, including workstreams, classification rule sets, queues, SLA definitions, and the decisions behind them, and the administrators are trained alongside the agents. A desk whose configuration is understood by one consultant who later leaves is the state most of the projects described on this page arrived in.
What we need from you: A decision on the cut over date at least two weeks out, agreement to close the old route on the day, and two administrators who will be trained to own the configuration.
Where this phase slips: Hypercare cut short to release the team onto the next thing. It saves a fortnight and costs a quarter, because the fixes that would have taken an hour during hypercare become change requests once the project has been closed.
What makes it eight weeks rather than twelve
Six conditions, and every one of them is a property of the project rather than something a partner can promise away. If all six are true, the plan above compresses to eight weeks because whole pieces of it disappear rather than because anybody works faster. If four are true, quote ten. If two are true, the honest answer is twelve and a conversation about which of them can be made true before week one.
One support team on one channel. Email plus one live channel is a release; email, chat, voice, SMS, and a customer portal at once is a programme, and it is the scope shape that most often becomes the open ended phase one described further up this page.
No migration beyond accounts, contacts, and open cases, taken from a system that can be read directly. That removes the second load, the second reconciliation, and most of weeks six to eight.
No custom development. The desk runs on configuration, which is the cheapest outcome available and the one we look for first on any build.
The planning decisions already answered and signed, so week one is environments and process observation rather than a discovery that ends in an argument.
A single named decision maker reachable inside the working day. This is worth more calendar time than any other item on the list.
Integrations limited to identity and at most one line of business system whose owner is genuinely on the project rather than being asked for favours.
Adding a second team, a second channel, or a customer portal afterwards repeats the sequence rather than stretching the first one, so the second release is shorter than the first and the third is shorter again. That is the argument for shipping one channel and letting the next one benefit from what it taught you, and it is the opposite of the everything at once first release that most often turns into the open ended phase one described further up this page.
What actually causes the overrun, and the week you find out
Implementations do not overrun because the estimate was wrong. They overrun because something that was true in week one was not visible until week eight, and the expensive property of each cause below is not what it is, it is how late it becomes apparent. The first three are the root cause families we test for on every takeover, written here in their timeline form. Each entry carries what it typically costs and the week it usually surfaces, so a project already running can be checked against it this afternoon.
Requirements nobody owns
Typical cost: Two to four weeks, and it is the one that recurs
Usually surfaces: Week 8 or 9, in testing, when the build meets the people it was built for
The build reflects what somebody assumed the business needed rather than how it works, so processes get configured that nobody follows and mandatory fields get filled with junk to get past a form. The cost is not the rework itself, it is that every increment after it is built on the same wrong picture, so the second estimate is as wrong as the first. This is the reason we ask for half a day beside real agents in week one rather than a requirements workshop, and it is why the pilot in week nine is given the authority to change the build instead of only to report on it.
Technical debt in the environment you are building into
Typical cost: Two to three weeks, more if there is no release path
Usually surfaces: Weeks 3 to 4, the first time a change has to be promoted
Almost no Customer Service build lands in an empty environment. It lands next to unmanaged layers stacked over managed ones, form JavaScript doing what a business rule should do, classic workflows still running beside the flows that were supposed to replace them, and plug-ins with logic buried in them and no tests. The visible symptom is that every change breaks something else, so the pace drops without anyone being able to say which week it dropped in. We inventory it in week two and quarantine what is load bearing rather than trying to fix it inside a delivery timeline, because a refactor of load bearing code in the middle of an implementation is how one project becomes two.
Modules configured wrong and then built around
Typical cost: One to three weeks, and it is the cheapest of the three to fix
Usually surfaces: Week 4, when routing is configured on top of it
The platform already does the thing the project is struggling with, but it was configured wrong or never configured at all, so somebody built around it: routing that never got past default queues, SLAs that pause and warn at the wrong moments, business process flows that do not match the stages the business uses, security roles cloned until nobody can say who sees what, duplicate detection switched on and never tuned. Remediation here is configuration rather than development, which is why it is worth finding in week four. Found in week ten it is not a configuration change any more, because two other things have been built on top of it.
Data nobody profiled before it was loaded
Typical cost: Two to four weeks, and this one has an external deadline
Usually surfaces: Weeks 6 to 8, at the first reconciliation
The load runs, the counts look plausible, and the numbers are wrong: created dates all clustered on the load date so every ageing report and SLA calculation built on that table is invalid, ownership sitting on the migration account so user and team scoped roles silently stop working, closed records that arrived open inflating the backlog, money columns that fell back to the base currency, and lookups that never resolved and were written as empty. Some of it is correctable in place. Records loaded without the created date override are not, because that value can only be set when the record is created, so they have to be reloaded, and reloading requires the source to still be readable. That is why the profiling happens in week two and the loads run twice.
Decision latency
Typical cost: Roughly a week for every decision that waits a fortnight
Usually surfaces: Continuously, and it is almost never on anybody's risk register
The single largest schedule risk on a healthy implementation is not technical. It is a question asked on Tuesday and answered at the following month's steering meeting, with the build either stopped or, worse, continuing on an assumption that gets reversed three weeks later. It is invisible in a status report because no individual delay looks like a delay. We manage it with one named decision maker, a fixed weekly checkpoint, and open questions written down with a name and a date against them rather than left to be settled by whoever reaches them first during the build.
Scope added mid build instead of displacing something
Typical cost: Three to four weeks per added channel or team
Usually surfaces: Weeks 3 to 6, when the build becomes visible enough to attract requests
Adding a channel, a portal, or a second team repeats the sequence rather than stretching it: its own routing, its own timeout behaviour, its own presence and capacity implications, its own training. The failure is not the request, it is accepting it without displacing anything of equal size from the current release. Change control with a displacement rule and a written roadmap date for what waits is the whole fix, and it is the difference between a twelve week project and the open ended phase one that eventually arrives on this page as a rescue.
No development, test, and production path, so test is production
Typical cost: One to two weeks now, and it is the debt every later change pays
Usually surfaces: Week 1 if anybody checks, week 10 if nobody does
The least interesting decision on the list and the one whose absence does the most damage. Without separate environments, work done in unmanaged solutions and shipped as managed ones, and the unpacked solution held in source control, every correction is typed into a live environment and nothing can be reviewed or rolled back. It costs a day at the start of week one. Retrofitted at week ten it costs a fortnight, and skipped entirely it produces exactly the layered, unreviewable configuration that makes a takeover necessary a year later.
Integrations to systems whose owners are not on the project
Typical cost: Two to three weeks, spent waiting rather than building
Usually surfaces: Weeks 6 to 7, when the interface is finally attempted
In this region the estate around Dynamics 365 routinely includes 1C, a local bank client, and national tax and e-invoicing services, and each one has an owner who was never invited to the kickoff. The delay is rarely the code. It is getting a test credential, an agreed message contract, and somebody on the other side who will answer during the same fortnight. Naming those owners in the week two integration inventory is a small piece of work that removes the most common cause of a slipped week seven.
The first three of those are the same three root cause families we test for on every assessment, and the remediation for each of them, together with a worked example of Customer Service unified routing that was never turned on, is set out in Dynamics 365 rescue services: how to fix a failed or stalled implementation. Why naming the cause has to come before any fix, and what happens to a project that treats the symptom instead, is the subject of name the root cause before you fix anything. If your project is still running and you want to know which of these eight is already in it before you commit to another date, that inspection is sold on its own as a fixed scope Dynamics 365 health check and technical audit, which is the cheapest thing on this page and the one most likely to save the timeline you already have.
When the timeline is already blown, this plan is the wrong question
Most people who search for how long an implementation takes are not planning one. They are checking a number they have already been given against something independent, usually because the last two dates did not hold. If that is you, the eight to twelve weeks above is not your answer, and neither is a fresh twelve week plan from anybody else. A project that has missed its date more than once has an undiagnosed cause, and any new estimate produced by the same process that produced the last one will miss too. The table below is the difference between the two situations.
The question
A healthy implementation
A project whose timeline is already blown
What the number actually is
Eight to twelve weeks from a signed scope to a live desk with hypercare finished, on the conditions listed above.
There is no number until the environment stops moving. A blown timeline cannot be re-estimated from inside itself, because the process that produced the first estimate is the thing that is broken.
What week one is spent on
The planning decisions, the environments and the release path, and half a day beside the agents who work the queue.
Environment lockdown, triage of what is blocking work today, and a forensic audit prioritized by what is bleeding. Nothing is estimated in week one because nothing is stable enough to estimate.
When users first see something working
Week nine, when the pilot team starts working real traffic on the real build.
Week four, and deliberately earlier, because a team that has already lost confidence in this project will not wait nine weeks to be shown something again.
What sets the end date
A scope agreed and signed in week two, with a displacement rule protecting it afterwards.
A fixed scope reset phase with a fixed price and a fixed end date at week six, ending in a go or stop decision that is yours to make with the audit and the debt inventory in hand.
How the data is treated
Profiled in week two, loaded twice into test, reconciled per table and signed by a business owner before cut over.
Already loaded and often already wrong. Every finding is classified as recoverable by reload, correctable in place, or genuinely lost, and the third list is written down and shown to the business rather than quietly absorbed.
What the timeline risk is
Decision latency and scope added without displacement. Both are managed by process rather than by effort.
The restore window, the date the source system goes read only, and how much undocumented knowledge walks out with the next person to resign. All three expire whether or not anyone is watching them.
What an honest partner says
Twelve weeks, and here are the six conditions that make it eight and the eight things that make it sixteen.
Six weeks to a stable environment, two working modules, and a written plan, then the remaining inventory as a dated backlog. Anyone quoting a fresh twelve week plan on a project that has already missed three dates is quoting the same plan that missed them.
The dated version of the right hand column is the fixed scope reset phase further up this page: six weeks, a fixed price, a fixed end date, and a named deliverable at every milestone, with the Environment Audit Report at week two and a working module in front of real users by week four. Where it is already an emergency rather than a slippage, the first seventy two hours run before any plan is written at all, because a timeline agreed on an environment that is still moving is a timeline that gets rewritten.
Customer Service & Omnichannel Implementation Guide
Customer Service is the module we are asked to rescue more often than any other, and it almost always fails for the same reason: it was configured in the wrong order, so decisions that had to be settled before the build got settled by whoever reached them first. Most implementation guides on this subject describe the phases of a project without ever naming a decision, which is precisely the advice that produces the projects described further up this page. This one is written the other way round: the decisions first, then the deployment sequence they feed, then the four areas that decide whether the result survives contact with real agents.
It is written to be usable without us. Each of the four areas below closes with the call we actually get when that part goes wrong, because the honest link between an implementation guide and a rescue service is that the second one exists for the implementations the first one did not save. If you are already past that point, the emergency rescue and stabilization process above is the right place to start instead.
The eight decisions to settle before anyone configures anything
These are expensive to change later and cheap to decide now. Every one of them appears in our audits as a finding on projects where it was never explicitly decided, which is not the same as being decided badly. An unmade decision gets made anyway, by whoever hits it first, and nobody writes it down.
1
Whether Dynamics 365 Customer Service is the right desk at all
The cheapest rescue is the one that never becomes necessary, and the first decision is whether this platform is the correct answer for the desk you are running. A team of four handling a steady shared mailbox does not need queues and entitlements. A team whose agents need the account, the contract, the last order, and the previous five cases in front of them while they answer does, and no standalone help desk gives them that without an integration to maintain. We set out the comparison honestly, including the cases where a different tool wins, on our Customer Service against Zendesk and Freshdesk page. Deciding this in week one is normal. Discovering it in month six is one of the ways projects arrive at this page.
2
What the unit of work actually is: a case, or a conversation
Everything downstream hangs off this. A case is a record with a lifecycle, an SLA clock, a resolution, and a reporting history. A conversation is a live session that may or may not deserve to become one. Decide which inbound traffic creates a case, which is handled and closed inside the conversation, and what triggers a promotion from one to the other, because SLA, entitlement consumption, and every duration metric you will later be asked for are properties of the case. Teams that leave this open end up with two parallel and irreconcilable pictures of the same workload, and neither number survives being questioned in a board meeting.
3
The queue topology, written as teams that exist today
Queues should mirror how the support organization is genuinely arranged, with one named human owner per queue, and they should be drawn before a single routing rule is written. An elaborate rule set feeding queues nobody owns is the most common routing failure we inherit, and it is not a routing failure at all. Put the list of queues, their owners, and the fallback queue on one page and get it signed. If that page cannot be written because nobody agrees who owns what, that is a finding about the organization rather than about the platform, and it has to be resolved before configuration rather than encoded into it.
4
What the SLA clock measures, and when it is allowed to stop
First response and resolution mean nothing until you have agreed the calendar they run against, whether the clock pauses while a case sits with the customer or with a third party, and what happens at warning and at failure. Write the pause conditions down in words a service manager would defend to a customer, then configure the SLA KPIs to match. This is the single decision most often skipped, because it looks like configuration and is actually a commercial commitment. A desk that reports resolution times computed on a clock nobody agreed will have its numbers rejected the first time they are inconvenient.
5
Entitlements: what each customer actually bought
If different customers are owed different response times, that difference has to live in the system rather than in the heads of two long serving agents. Entitlements record the terms, the number of cases or hours covered, the channels included, and which SLA applies, and they decide what an agent sees on the case before they promise anything. Decide in planning whether entitlements are in scope for the first release. They frequently should not be, and saying so explicitly is better than half building them, which is the state we find them in more often than any other.
6
Who is allowed to see which case
Case visibility is a data model decision disguised as an admin task. Business units, teams, and the ownership model determine what is possible, and they are expensive to change once there are hundreds of thousands of rows and a dozen integrations resolving records by owner. Decide early whether agents see everything, only their queue, or only their business unit, and whether any category of case has to be genuinely private. Security roles cloned until nobody can say who sees what is one of the three most common findings in our audits, and it starts here.
7
Channel scope for the first release, and what waits
Email plus one live channel is a first release. Email, chat, voice, SMS, and a customer portal at once is a programme, and it is the shape of scope that most often becomes the open ended phase one described further up this page. Each channel carries its own routing, its own timeout behaviour, its own presence and capacity implications, and its own agent training. Pick the channel that carries the most volume, ship it, and let the second channel benefit from what the first one taught you. The channels that wait should be written on the roadmap with a date, because scope that is deferred without a date comes back as a change request in week three.
8
Environments, solutions, and the release path, before anything is built
Separate development, test, and production environments, work done in unmanaged solutions and shipped as managed ones, and the unpacked solution held in source control. This is the least interesting decision on the list and the one whose absence does the most damage, because without it every later correction is typed into a live environment and the system accumulates the layered, unreviewable configuration that makes a rescue necessary. Settle it in week one. It costs a day at the start and it is what allows every subsequent change to be reviewed and rolled back.
The first of those decisions is the one people expect a Dynamics 365 partner to answer dishonestly, so we answer it in public: the comparison against the standalone help desks, including the cases where one of them is the better buy and the cases where the Microsoft 365 tools you already own are enough, is set out in Dynamics 365 Customer Service against Zendesk and Freshdesk.
The deployment sequence, and what proves each step is done
The windows below are what a single team Customer Service rollout normally takes when the decisions above have been made. The order is not a preference. Each step depends on the one before it being genuinely finished, and the most common way an implementation goes wrong is that two of these ran in parallel and neither was ever proven. Where a channel, a portal, or a second team is added, the sequence repeats rather than stretches.
Week 1
Settle the decisions, then write them on one page
Before any configuration, the eight decisions above are answered and recorded: the unit of work, the queue topology and its owners, what the SLA clock measures and when it pauses, whether entitlements are in scope, the visibility model, the channels in the first release, and the environment and release path. The page is short on purpose, because its job is to be argued with now rather than discovered later. Where a decision genuinely cannot be made yet, it is written down as an open question with a name and a date against it, not left to be settled by whoever reaches it first during the build.
Proof it is done: A signed one page decision record, and the environments and solution structure in place.
Week 1 to 2
Model the case before configuring anything that reads it
The case table, its relationship to contact and account, the option sets that will later carry your reporting, parent and child case behaviour, and the columns your routing will test. Case reason, subject, and channel are governed lists rather than free text from the start, because a report grouped over free text is not a report. This is also where the small number of custom columns get justified one at a time, since the difference between a Customer Service build that stays supportable and one that does not is largely how much was added to the case form and why.
Proof it is done: The case model and its option sets, deployed as a managed solution into the test environment.
Week 2
Get intake working before routing exists
The mailbox is approved and enabled for server side synchronization and confirmed as actually receiving, then automatic record creation rules turn inbound mail into cases with the same handling on every channel. Intake is deliberately proven before any routing rule is written, because a mailbox that quietly failed test and enable stops email routing before a single rule runs, and a team that has not separated the two will spend a week debugging rules that were never reached.
Proof it is done: Real inbound mail creating correctly shaped cases in the test environment.
Week 3 to 4
Unified routing, classification first
The workstream with its type, distribution mode, and capacity model, then the work classification rule sets that stamp the attributes routing will read, then the route to queues decision list evaluated top to bottom with narrow rules above broad ones and a catch-all at the bottom pointing at a fallback queue somebody watches. Queues are created and given their members and their assignment method last. The order matters more than any individual rule: classification stamps, the decision list acts, and routing that looks correct on screen but never fires is almost always an attribute that was never set.
Proof it is done: Routing diagnostics on real traffic showing which rule matched and how assignment resolved.
Week 4
SLAs, business hours, and entitlements if they are in scope
The customer service calendar and business hours first, because every duration is computed against them, then the SLA KPIs for first response and resolution with the pause behaviour that was agreed in writing in week one, then warning and failure actions that reach a person rather than only writing a value. Entitlements are configured here or explicitly deferred with a date. Each SLA is tested against a deliberately awkward case that crosses a weekend and sits with the customer for two days, because that is the case a service manager will be asked about.
Proof it is done: SLA behaviour demonstrated on a test case that crosses non working hours and a customer pause.
Week 5
Knowledge, seeded from what the desk actually answers
Drafting, review, versioning, and expiry are enabled from the start, and the first articles are written against the top case reasons in the last quarter of real traffic rather than from a workshop wish list. Search is proven from inside the case form and the conversation panel using the phrases agents genuinely type, and articles are linked to the cases that used them so deflection becomes measurable instead of asserted. Every article gets an owner and a review date on a named person, because an article with neither becomes wrong quietly and takes the agents trust in the whole knowledge base with it.
Proof it is done: A published starter set covering the top case reasons, each with an owner and a review date.
Week 5 to 6
The agent workspace and the live channel
Session and workspace behaviour, what closes after wrap up, what survives a refresh, presence and capacity profiles, notification timeouts, and the customer inactivity warning and closure values on the workstream. Every one of these is checked against how long a real conversation in your business takes, including its quiet parts, rather than left at the default. Where a live chat channel is in scope, the reconnection window is configured so a customer who loses the page rejoins the same conversation instead of starting again with a different agent.
Proof it is done: Timeout, presence, and reconnection settings tested against a deliberately slow conversation.
Week 6 to 7
Pilot with a real team on real traffic
A named pilot group works live traffic with the authority to change the build before it reaches everybody else, and the routing diagnostics, SLA outcomes, and knowledge search results are read daily rather than at the end. This is the point at which the queue topology decided in week one meets the people who work it, and it is normal for it to change here. A pilot whose feedback cannot alter the configuration is not a pilot, it is a rehearsal, and it is the reason so many desks go live with problems everybody already knew about.
Proof it is done: A written pilot log of what changed as a result, and the go or hold decision.
Week 7 to 8
Baseline, cut over, and hypercare
The current numbers are captured before go live so the improvement claim later has something to compare against, then the desk cuts over with a named person answering agent questions in the channel agents already use and a written daily note of what changed. Hypercare runs for two weeks minimum. The build is documented as it actually is, including the workstreams, rule sets, queues, SLA definitions, and the decisions behind them, because undocumented routing drifts back into trouble the first time the team reorganizes.
Proof it is done: A pre go live metric baseline, the handover pack, and two weeks of daily hypercare notes.
The four areas that decide whether it works: routing, knowledge, reporting, adoption
A Customer Service implementation is rarely lost on the case model. It is lost in one of these four, and each one has its own planning decisions, its own best practice, and its own recognizable failure. The failure is stated at the end of each because it is what makes the rest of it checkable: if you already recognize the call we describe, that part of your desk is the part to look at first.
Routing rules
Unified routing is configuration rather than development, and it fails in a small number of very repeatable ways. Almost every broken routing setup we inherit is broken for one of three reasons: the attribute the decision list tests was never stamped, there is no catch-all at the bottom, or the queues were designed before anyone asked who owns them.
Decide before you configure
Push or pick per workstream, and the capacity model, because the distribution mode and the capacity model are what everything else hangs off.
Whether the routed unit is a record such as a case, or a live conversation, since that decides where the SLA and the reporting live.
The attribute every routing decision reads, such as topic, product, language, or contract tier, and where in the process it gets stamped.
Whether skill based routing is real for you. Skills only work if somebody maintains the skills model as people join, leave, and get trained.
Which queue is the fallback, and which human being reads it. Work that matches nothing goes there, and it is where unowned work goes to be forgotten.
How we build it
Prove intake before writing a rule. A mailbox that failed test and enable stops email routing before any rule is evaluated.
Classification rule sets first, decision list second, narrow rules above broad ones, and a catch-all at the bottom pointing at a queue with an owner.
Keyword and condition rules as the opening move. Machine learning classification earns its place where the language is genuinely ambiguous, and it is a poor first step because a model layered over queues that are already wrong makes misrouting harder to explain.
Where the routing attribute depends on something outside the message, such as the contract a customer is on, derive it in a flow or a plug-in, write it to a column, and route on that column.
Read the diagnostics on real traffic, which show which classification rule matched, which decision list rule chose the queue, and how assignment resolved. That is what separates an unmatched condition from a mailbox problem from an agent capacity problem.
Document the workstreams, rule sets, and queues afterwards, because routing drifts the first time the team reorganizes and nobody remembers why rule seven exists.
When it goes wrong, this is the call we get
The call we get: everything lands in one queue, a team lead triages it by hand every morning, and the conclusion inside the business is that the platform cannot route by topic. It can, and in most cases the remediation is configuration measured in days rather than a rebuild. This is the most common and the cheapest of the three root causes we test for on a rescue.
Knowledge management
A knowledge base is the only part of a Customer Service implementation that decays on its own. Everything else stays as configured until somebody changes it. Articles go out of date whether or not anybody touches them, so the governance decisions matter more than the tooling ones.
Decide before you configure
Who writes, who approves, and who retires. Three roles, named people, agreed before the first article rather than after the hundredth.
Internal only for the first release, or external from the start. This drives the tone, the language, and whether a customer portal is in scope at all.
The taxonomy, built from the words agents and customers actually use rather than from the organization chart.
Whether articles are linked to the cases that used them, which is the only way deflection and article quality ever become measurable.
The review cadence and what happens to an article that misses it. An article with no expiry date is a promise nobody renewed.
How we build it
Seed the base from the top case reasons in the last quarter of real case data, not from a workshop list of what the team thinks it gets asked.
Turn drafting, review, versioning, and expiry on from day one. Retrofitting governance onto four hundred existing articles is a project of its own.
Make search work from inside the case form and the conversation panel, tested with the phrases agents genuinely type, so nobody has to leave the record to find an answer.
Link the article to the case at the point of use, then review article performance monthly and retire what nothing links to.
Give every article an owner and a review date that lands on a named person, because the fastest way to lose agent trust in a knowledge base is one confidently wrong article.
When it goes wrong, this is the call we get
The call we get: a knowledge base of several hundred articles, most of them drafts nobody approved, a search that returns nothing for the phrases agents type, and a support team that has quietly gone back to asking the two people who know. The content is usually salvageable. The governance never existed.
Reporting
Most support reporting problems are not reporting problems. They are modelling decisions taken months earlier, because the numbers a service manager will be asked for are stamped at case creation and at each transition, and they cannot be reconstructed afterwards from data nobody wrote.
Decide before you configure
The five to eight numbers the service manager will actually be held to, agreed before the case model is built rather than after go live.
Whether resolution means a resolved case or a resolved customer problem. Reopen rate is the honest test, and it only exists if you decided to measure it.
The business hours and calendar every duration is computed against, which is the same decision as the SLA pause behaviour and should be made once.
Which numbers come from the built-in historical analytics and which need Power BI over Dataverse. Both read the same data, and choosing per metric avoids building a warehouse for four numbers.
Whether conversation metrics and case metrics are reported separately, because averaging a two minute chat with a four day case produces a number that describes nothing.
How we build it
Configure first response and resolution as SLA KPIs with agreed pause behaviour, rather than approximating them with a due date column and a scheduled reminder.
Keep case reason, subject, and channel as governed option sets. A Pareto over free text is not a Pareto, and it is the most common reason a support report cannot answer why volume went up.
Capture a baseline in the two weeks before go live, so the improvement claim afterwards has something to compare against rather than a recollection.
Report backlog age alongside volume. Volume alone hides a queue that is quietly growing older, which is the metric that predicts the next crisis.
Put the reports in front of the people who will be measured by them during the pilot, because a metric that is disputed after go live gets abandoned rather than fixed.
When it goes wrong, this is the call we get
The call we get: a desk that has been live for a year and cannot answer first response time by customer, because no path in the build ever stamped it. Fixing the reporting means fixing the model, and the history that was never recorded does not come back. This is the argument for settling the metrics before the case model rather than after.
Agent adoption
Adoption is the pillar that decides whether the other three mattered, and it is the one most often treated as a training slot in the final week. Agents do not abandon a system because it lacks features. They abandon it because it costs them time on every single case, or because it drops them mid conversation.
Decide before you configure
Who the pilot team is, and whether they have real authority to change the build before it reaches everyone else.
The session and workspace behaviour: what closes after wrap up, what survives a refresh, and how many sessions an agent is expected to hold at once.
Every timeout value, checked against how long a real conversation in your business takes including its quiet parts, rather than left at a default.
Who answers an agent question on day one, and in which channel. If the answer is a ticket queue, adoption is already in trouble.
The language training and documentation are delivered in, which for a team in this region is frequently not English.
How we build it
Put agents in the room during configuration rather than only at user acceptance testing. The people who work the queue know why the queue topology is wrong, and they know it in week two rather than week eight.
Design the case form for the first ninety seconds of a call. What the agent needs to see and set immediately goes on the first tab, and everything else moves off it.
Settle the environment session and inactivity timeouts, the conditional access sign-in frequency, presence and capacity, and the browser policy on the agent fleet before go live, not after the first week of reported disconnections.
Make single tab working an explicit rule in the agent guidance, because presence is per user rather than per tab and a second session interferes with the first.
Run at least two weeks of hypercare with a named person, in the channel agents already use, with a written daily note of what changed so nobody has to ask twice.
Train the administrators as well as the agents. A desk whose configuration only one departed consultant understood is the state most of the projects on this page arrived in.
When it goes wrong, this is the call we get
The call we get: the build is technically correct and nobody uses it. Agents are being signed out mid chat, the form takes twelve clicks to close a case, and the shared mailbox everybody was told to stop using is quietly back in service. Every one of those is fixable, and none of them is fixed by more training.
The routing configuration itself is written out step by step, from enabling unified routing through workstreams, classification rule sets, the route to queues decision list, and reading the diagnostics afterwards, in name the root cause before you fix anything, which uses routing that was never turned on as its worked example of the most common and most fixable root cause. The agent adoption settings, including the environment session and inactivity timeouts, conditional access sign-in frequency, workstream inactivity rules, presence, and the browser policies that quietly tear down a live conversation, are diagnosed one by one in fixing Omnichannel agent disconnections and chat timeouts.
For what this looks like when it goes to plan at scale, our CRM modernization for a European bank put Sales and Customer Service live on a single Dataverse model with the onboarding process automated in Power Automate, delivered by three senior consultants in five months. The relevant part for a Customer Service implementation is not the headline number, it is that a service desk and a sales team shared one customer record from the first release rather than being integrated to each other afterwards.
If the implementation has already gone wrong
This guide is what we do when we are there from the start. The rest of this page is what we do when we are not. A Customer Service or Omnichannel rollout that has stalled, gone live and been abandoned by its agents, or been left half configured by a partner who is no longer answering does not need the sequence above run again from week one. It needs the environment locked down, the root cause named, and the parts that are sound kept. In most cases the remediation turns out to be configuration rather than development, which is the cheapest outcome available and the one we look for first. We take that work on directly, or on a white-label basis for other Microsoft partners.
Rescuing Field Service Resource Scheduling Optimization
Resource Scheduling Optimization only produces schedules the dispatch team can trust when the inputs it depends on are correct. When a Field Service rollout is in trouble, these are the inputs we audit and put right so the engine sends the right technician to the right job.
Skills and characteristics that do not match the work
Resource Scheduling Optimization can only assign a technician to a job when the resource carries the characteristics the work order requirement asks for, at the proficiency the requirement expects. When skills are missing, inconsistently named, or set at the wrong rating, the engine either cannot fill the booking or fills it with the wrong person. We audit the characteristic model against the real requirements and put the skills matrix back into a state the optimizer can use.
Locations and geocoding the engine cannot trust
Optimization depends on knowing where work is and where resources start, travel from, and return to. Missing or wrong geocodes on accounts, work orders, and resources, or a start and end location that does not reflect reality, produce impossible travel times and schedules the dispatch team throws away. We fix the location and geocoding data and the resource start and end settings so travel and proximity are calculated against the truth.
Requirements, territories, and scheduling parameters
Resource requirements, bookable resources, territories, working hours, and the scheduling parameters and optimization goals all have to line up before an automated schedule makes sense. We review how requirements are generated, how resources and territories are set up, and how the optimization scope and goals are configured, and we correct the mismatches that make the engine behave in ways the business did not intend.
Documentation so the schedule stays trustworthy
A tuned Field Service and RSO setup drifts back into trouble when nobody records how it was configured. We document the characteristics, requirement templates, territories, scheduling parameters, and optimization goals so your dispatchers and administrators can keep the schedule accurate as skills, resources, and territories change over time.
Frequently Asked Questions
What is a Dynamics 365 project rescue or takeover service?
It is a service where a new partner takes over a stalled or failed Microsoft Dynamics 365 or Power Platform implementation and gets it to a working state. Solzet begins with a forensic audit to diagnose the root causes, stabilizes the environment so daily work and data are safe, agrees a prioritized recovery plan, and then delivers the fixes and completes the build. The aim is to save the investment already made and keep what is sound, rather than throw everything away and start again, unless a review shows the current design genuinely cannot be salvaged.
When should we bring in a rescue partner instead of pushing on?
The usual signals are a project that is over budget or past its deadline with no credible end date, an original partner who has stalled or walked away, users who have lost confidence and gone back to spreadsheets, an environment that is unstable or full of errors, and a build that nobody can explain because it is undocumented. In Field Service specifically, it is time when Resource Scheduling Optimization returns schedules the dispatch team cannot trust and the board is still worked by hand. Any one of these on its own is a reason to get an independent diagnosis before spending more.
What happens in the first 72 hours of an emergency intervention?
Three things run at once. First, environment lockdown: we freeze deployments into production, cut System Customizer and System Administrator roles back to the people who need them, stop edits being made directly in the live environment, copy production into a sandbox so the failing state is preserved, turn off the flows and jobs that are actively damaging data, and confirm what restore points exist before that window closes. Second, emergency triage of the failures blocking daily work, reproduced and fixed in the sandbox and promoted through a controlled path rather than typed into production. Third, a forensic audit prioritized to answer what is causing harm, what will fail next, and what must not be touched yet. Alongside all of it we put a communication plan in place: one named decision maker, one channel of record, and a daily written brief. The week closes with a written stabilization statement that is yours whatever you decide next.
Can you take over from our current vendor mid project, and how fast can that happen?
Yes, and the handover is usually faster than clients expect, because a Dynamics 365 environment is tenant property rather than partner property. Whoever owns the Microsoft 365 tenant can recover Global Administrator, reassign Power Platform Administrator and Dataverse System Administrator, take ownership of the environments and of the application users and service principals the integrations authenticate as, and claim the repositories and deployment connections holding the solution source, regardless of who built it and regardless of how the relationship ended. On an emergency vendor takeover we move tenant administration in the first four hours, Power Platform and Dataverse administration within eight, the integration identities within a day, and source and pipelines within a day. Licensing and any question about ownership of custom code run alongside rather than on the critical path. We begin on read access where full rights take longer to arrange, because waiting for permissions is the most common way an emergency loses its first day.
What do you need from us to start an emergency takeover within 24 hours?
Three things. Someone on your side who can authorize access to the environment, a named decision maker who can approve a freeze or a change during the working day rather than at a weekly steering meeting, and permission to stop things: deployments into production, direct edits in the live environment, and any flow, job, or integration found to be damaging data. Read access is enough to start the audit, so nothing waits on full administration rights. Everything else runs from our side once those three are in place: the vendor takeover of tenant and environment administration, the sandbox copy of production, the triage of the failures blocking daily work, and the first daily written brief.
What is the difference between stabilization and a standard consulting engagement?
A standard engagement books discovery workshops for a few weeks out, leaves the environment as it is in the meantime, reads documentation and interviews rather than the system itself, and produces a findings deck at the end of a discovery phase on a time and materials basis. Emergency stabilization locks the environment down on day one, works in a sandbox copy of production, reads the solution layers, code, flows, security roles, integrations, and run history directly, releases the first emergency fixes inside the first week, and reports daily in writing. The people who diagnose it are the people who fix it, the reset phase that follows carries a fixed price and a fixed end date, and the audit, debt inventory, and handover pack are written to be usable by any partner or by your own team.
Our Dynamics 365 system is broken and we have no budget for a rescue. What can we do this week?
Six things, all free, in this order. Secure administrative access first: confirm at least two people inside your own organization hold Global Administrator and Power Platform Administrator, and that someone other than the outgoing partner holds System Administrator in each environment, because the tenant and the data are yours regardless of who built the system. Freeze change with one written message, cut the System Customizer and System Administrator roles back to the people who need them, take a manual backup in the Power Platform admin center, and write down the date your oldest usable restore point expires. Document the customizations while people still remember them: export the unmanaged solutions and any flows living outside a solution, export the records the business cannot lose, and write a plain inventory of what was built and who last touched it. Disable the automations that are actively damaging data, turning them off rather than deleting them and keeping a list of what you disabled and why. Establish one manual fallback per broken process, mandated rather than invented team by team, with spreadsheet headings copied exactly from the fields in Dynamics 365 so the data can be loaded back later. Then write the state of the system on one dated page. None of this repairs anything, and all of it stops the situation getting worse and produces the evidence you need to get a repair funded.
How do we justify the cost of a Dynamics 365 rescue to a board that has already paid for a failed project?
By replacing adjectives with arithmetic. Measure six things you already hold the numbers for: the hours a week your manual workaround costs across every team at loaded cost, the work the system is no longer doing compared with a normal month before the crisis, the count of damaged records against the date your restore window closes, the licences still being billed every month for people working in spreadsheets, the knowledge that leaves with one or two named people, and any compliance or audit obligation the workaround breaks. Then present it correctly. State the loss as a cost per week rather than a total, so it keeps running while the decision is not made. Attach a deadline that did not come from you, such as the restore window closing or a departing administrator, because requests without dates are never approved. Ask for the smallest thing that unblocks the next decision, which is a bounded fixed price assessment rather than a rescue programme of unknown size. Price three options including doing nothing, with the weekly rate as the price of that option. And take it to the director whose team is doing the work by hand rather than to the IT budget that has already been spent on this project once.
Can we fix a failed Dynamics 365 implementation ourselves without a partner?
Partly, and the honest line is worth knowing before you start. Freezing change, taking backups and exports, documenting the build, disabling harmful automations, and running a controlled manual fallback are all safe to do alone and are genuinely valuable. Five situations are where continuing alone starts costing more than it saves: data that is already damaged rather than simply missing, because correcting records by hand without classifying them first is where DIY does the most harm; a production restore, which is all or nothing and destroys every legitimate record entered since the backup; a fault inside plug-ins, custom API, form scripts, or stacked solution layers, where no admin screen will show you the cause; a change freeze that cannot be lifted after a fortnight because there is no test environment or release path, at which point the freeze has become the failure; and a manual workaround that has quietly become the system. In all five the cheapest next step is a bounded independent assessment rather than either continuing alone or committing to a full rescue.
Our go live date has slipped three times. Is it too late for a rescue?
Almost never. A repeatedly missed go live is a symptom of an undiagnosed root cause rather than proof that the build is worthless, and in most cases a substantial part of the environment is worth keeping. What it does mean is that the situation has become an emergency, because by the third missed date the business has usually stopped preparing for a go live it no longer believes in and the delivery team is running on burnout. We intervene the same way regardless: lock the environment down, stop the damage, diagnose what is actually causing the slippage, and put a date on the board that has a named deliverable behind it. Where a review shows part of the design genuinely cannot be salvaged, we say so in writing and scope the smallest rebuild that reaches a working solution.
Will you have to rebuild everything from scratch?
Usually not. A takeover is not a restart. In most cases the goal is to save the investment already made, keep the parts of the build that are sound, and fix only what is actually broken. The forensic audit tells us what to keep, what to rework, and what, if anything, must be rebuilt, and the plan scopes the smallest rebuild that gets you to a working solution. Where the honest answer is that a component cannot be salvaged, we say so plainly rather than layer more work on a broken foundation.
Can you fix Dynamics 365 Field Service Resource Scheduling Optimization?
Yes. Resource Scheduling Optimization only produces schedules the dispatch team can trust when the inputs it relies on are correct: technician skills and characteristics at the right proficiency, resource and requirement records, territories, working hours, accurate locations and geocoding, and sensible scheduling parameters and optimization goals. We audit those inputs against the real work, correct the mismatches that send the wrong technician to the wrong job, and document the configuration so the schedule stays accurate as skills, resources, and territories change.
What is the fixed scope reset phase and how long does it take?
It is the first six weeks of a Solzet takeover, sold at a fixed price with a fixed end date and a named deliverable at every milestone. Days 1 to 5 are access and triage of the failures hurting users now, ending in a triage log, an immediate risk list, and the first emergency fixes released. Weeks 1 to 2 produce the Environment Audit Report and the Technical Debt Inventory. Week 2 is stakeholder realignment, ending in a one page Reset Scope with the price, the end date, and the acceptance criteria. Weeks 3 to 4 deliver the first working module to real users. Weeks 5 to 6 deliver a second increment, a clean development, test, and production release path, a written handover pack, and a go or stop decision that is yours to make.
How soon will we see visible progress on a rescue?
Within the first week. Triage runs alongside the audit rather than after it, so the failures that block daily work, such as a flow overwriting records, a form that will not save, or a schedule sending crews to the wrong place, are reproduced in a sandbox, fixed, and released inside days. The first documents land at the end of week two and the first working module is in front of users by the end of week four. We time box it this way on purpose, because most rescue proposals offer an undated assess, stabilize, and optimize sequence with nothing a user can see for months, and a team that has already lost confidence in its Dynamics 365 project will not wait that long.
What is in the technical debt inventory?
Every finding from the code and configuration audit, recorded as a line with what it is, what it blocks or breaks today, the risk of leaving it through the next release wave, and the rough effort to fix it. Each line is marked keep, rework, rebuild, or delete, and the customizations are split into the ones that are load bearing, meaning the business genuinely runs on them, and the ones that are pure risk. It covers solution layering, plug-ins and custom APIs, form JavaScript, Power Automate flows and classic workflows, the Dataverse model, security roles, integrations, and PCF or canvas components. The inventory is yours and is written to be actionable by any partner or by your own team.
How do you fix a Dynamics 365 project that has lost control of its scope?
In three phases. First the backlog is frozen so nothing new enters while it is counted, with a single parked list for anything urgent that arrives during the freeze. Then every item is marked in scope, added later, or a reworded duplicate of something already delivered, traced to a decision the business actually takes, and re-baselined into a one page Reset Scope of ship now, ship later, or drop, with the drop list shown to the people who asked for those items rather than deleted quietly. Then change control reopens with one named decision maker and a displacement rule: a new item of a given size removes something of equal size from the current release instead of adding to it. We also read the environment for scope that was built and never used, such as apps with no security role assigned and tables holding a handful of rows, because that is scope already paid for that costs testing and maintenance on every release and can be cut for nothing.
Can a bad Dynamics 365 data migration be fixed after go live?
Most of it, but the categories are not equal and one of them has a deadline. We first stop any delta or incremental load that is still running, because a migration that is still writing will undo corrections as fast as they are made, and we confirm how long the source system stays readable, since that date sets the limit on anything that has to be re-extracted. Every finding is then classified as recoverable by reload, correctable in place, or genuinely lost. Records loaded without the created date override fall into the first group and have to be reloaded, because that value can only be set when a record is created. Ownership, state and status, currency, and unresolved lookups are usually correctable in place. Corrections run in dependency order, one pass at a time, in a sandbox first, and the work closes with a per table reconciliation the business signs plus alternate keys and duplicate detection rules so the same damage cannot return through the same door.
How do you reduce customization debt without breaking what the business runs on?
By fixing the release path before touching any code, then working from the safest changes to the riskiest. Separate development, test, and production environments come first, with work done in unmanaged solutions and shipped as managed ones and the unpacked solution in source control, so any change can be reviewed and rolled back. Then the pure risk goes: unreferenced scripts, disabled workflows, half finished components, and code doing what business rules and calculated or rollup columns already do, each removed after a dependency check and released on its own. Then every column that has more than one writer gets exactly one owner, and classic workflows are retired deliberately into a flow or a plug-in rather than left running beside their replacement, which is what removes the intermittent bug nobody can reproduce. The customizations the business genuinely depends on are rewritten last, one component per release, behind the behaviour they already have. A single big refactor of load bearing code is how a rescue becomes the next failed project.
Our integrations seem fine but data keeps going missing. How do you find the cause?
By assuming the failures are invisible rather than absent. For each interface we ask one question: if this failed at three in the morning, who would find out and how. A flow with no failure path, a plug-in that swallows its exception, or a queue with no dead letter handling all fail silently by design, so we start by counting what has already failed in the last month of run history, which on a stalled project is usually in the hundreds. We check what each interface authenticates as, since a connection owned by a person who has left breaks the day their account does, and we check the errors for service protection throttling before rewriting any logic, because the answer there is retry with backoff and batching rather than a redesign. Before recovering anything we make replay safe with alternate keys and upsert behaviour, since replaying into an interface that creates rather than updates turns data loss into duplication. The backlog is then replayed and reconciled in both directions, and a scheduled reconciliation with an alert on divergence is left behind so the next failure is visible in a day rather than a quarter.
What does the forensic audit actually look at?
We examine the real solution rather than the story around it: the Dataverse data model, forms and views, business rules, plug-ins and custom code, Power Automate flows, security roles, integrations, and any PCF or canvas components, compared against what the business needs. For Field Service we also look at requirements, resources, territories, locations, and the optimization configuration. The output is a written diagnosis of what is sound, what is broken, what is missing, and which problems are causing the symptoms you see, which becomes the basis for the recovery plan.
Can you recover a failed Dynamics 365 project in Armenia?
Yes, and it is work we do from Yerevan rather than remotely into Armenia. Solzet is an Armenian Dynamics 365 Customer Engagement and Power Platform consultancy, so a recovery here runs on the same clock, the same holiday calendar, and the same law as your own company, and an engineer can be at your office in Yerevan the same morning. The method is the one set out on this page: environment lockdown, triage of what is blocking work today, a forensic audit under emergency conditions, and a fixed scope reset phase with a fixed price and a fixed end date. What is specific to Armenia is the 72 hour local path through it: confirming whose tenant your data is actually in, because on smaller local projects it is sometimes still inside the contractor's own subscription, freezing the interfaces to 1C, accounting systems, the bank client, and the State Revenue Committee portal before anything cosmetic, and putting every verbally agreed decision into a written brief the same evening.
How fast can you get on site in Yerevan for an emergency Dynamics 365 takeover?
An engineer can be at any office in Yerevan inside the hour during the working day, and Gyumri, Vanadzor, and most regional centres are a half day by road. That matters on the first day of a rescue more than at any other point, because the things that stall a lockdown are people problems rather than technical ones: an administrator password nobody wrote down, the single laptop holding a working connection to the accounting system, and the colleague who will demonstrate the real process but will not write it down. On site attendance here is an ordinary working decision rather than a travel budget, so it can happen again later in the rescue instead of being spent once at kickoff.
The company that built our system is a local Yerevan firm and the relationship has broken down. What now?
It is a more workable position than it sounds. First, establish whose tenant the environments are in. If the tenant is yours, you can recover Global Administrator, reassign Power Platform and Dataverse administration, and take ownership of the environments and the identities the integrations run as, regardless of how the relationship ended and without the other side agreeing to anything. If the solution was built inside the contractor's own subscription, which does happen on smaller local projects, treat it as a migration rather than a takeover and secure a full export of your data and the unpacked solution while access still exists. Second, ask for one hour with the people who built it. The Armenian market is small enough that they are usually still reachable, and that hour recovers what no audit can. We keep that conversation about the system rather than about fault, which is both how the answers get given and simple realism about a market where everyone works together again.
Is there a Dynamics 365 takeover specialist in a similar time zone to the CIS and Caucasus?
Yes. Solzet is a Dynamics 365 Customer Engagement and Power Platform consultancy in Yerevan, Armenia, which sits in GMT+4 and does not observe daylight saving, so the offset never moves through the year. That is the same clock as Tbilisi and Baku, one hour ahead of Minsk and Moscow, one hour behind Tashkent and Almaty, and two behind Bishkek. For a business anywhere in that range a takeover runs inside a single shared working day: the daily written brief lands in your morning, decisions are made in an afternoon call rather than an overnight email cycle, and emergency stabilization can actually happen on the day it is needed. We work in Russian, Armenian, and English, and Yerevan is about an hour from Tbilisi and under three from Moscow, so on site attendance is affordable at short notice rather than a planned event.
Why use a regional partner instead of a global consultancy for a project takeover?
For a multi country programme with heavy governance obligations, a global firm such as Avanade is often the right answer, and their Dynamics 365 practices are genuinely strong. A takeover is a different job. It rewards live overlap with your working day, the ability to read the project in the language it was specified in, senior engineers who diagnose and then fix the same system rather than handing a report to a delivery bench, and on site presence that does not carry long haul travel time and per diems on top of an already high rate. A Yerevan team gives you those four things structurally rather than as a commitment, and it reads the estate around Dynamics 365 as it actually is in this region, including integrations to 1C, local bank clients, and national tax and e-invoicing services.
What happens if the previous partner will not hand over access or documentation?
You are in a better position than you probably think, because a Dynamics 365 environment is tenant property. Whoever owns the Microsoft 365 tenant can take back Global Administrator, reassign Power Platform Administrator and Dataverse System Administrator roles, and take ownership of the environments, the application users, and the service principals the integrations run under, regardless of who built the solution. We start a takeover with a written access list covering exactly that, plus the repositories where the solution source lives and the deployment service connections. Where there is genuinely no documentation and no one left to ask, the audit reconstructs the design from the environment itself, which is slower but is a normal part of the work rather than a blocker.
Are you a Dynamics 365 Customer Service implementation partner as well as a rescue partner?
Yes, and the two are the same skill applied at different points. Solzet implements Dynamics 365 Customer Service and Omnichannel from scratch, and we are also the team that gets called when someone else's implementation stalls, which is what makes the implementation advice worth reading: we see the failure modes from the inside every month. Our Customer Service and Omnichannel implementation guide on this page sets out the eight decisions to settle before configuration, the deployment sequence with the proof that closes each step, and the four areas that decide the outcome, which are routing rules, knowledge management, reporting, and agent adoption. We deliver directly and on a white-label basis for other Microsoft partners.
What are the key planning decisions in a Dynamics 365 Customer Service implementation?
Eight, and all of them are cheap now and expensive later. Whether this platform is the right desk for you at all. Whether the unit of work is a case or a live conversation, since SLA and reporting are properties of the case. The queue topology, drawn as teams that exist today with one named owner per queue and a fallback queue somebody genuinely reads. What the SLA clock measures and when it is allowed to pause, which is a commercial commitment rather than a configuration setting. Whether entitlements are in scope for the first release or explicitly deferred. Who is allowed to see which case, which is a data model decision disguised as an admin task. The channel scope of the first release, because email plus one live channel is a release and five channels at once is a programme. And the environment, solution, and release path, which is the least interesting decision and the one whose absence does the most damage.
How long does a Dynamics 365 Customer Service implementation take?
Eight to twelve weeks for a single support team, measured from a signed scope to a live desk with hypercare finished. Weeks 1 to 2 are discovery: the planning decisions, the environments and the release path, the case model and its governed option sets, and a profile of the data and the integrations rather than a description of them. Weeks 3 to 5 are configuration, in a fixed order: intake proven before any routing rule exists, then unified routing with classification first, then business hours, SLA KPIs, entitlements or an explicit deferral, and knowledge seeded from real case reasons. Weeks 5 to 8 are the agent workspace and live channel, the small amount of custom code that survived justification, the integrations, and two full data loads each closed by a reconciliation. Weeks 8 to 10 are testing on deliberately awkward cases and then a pilot on real traffic with the authority to change the build. Weeks 10 to 12 are the pre go live baseline, cut over, and at least two weeks of hypercare with documentation and administrator training. Eight weeks is one team on one channel with no migration and no custom development. Twelve weeks is the same desk with a real data load and integrations to systems somebody else owns.
What makes a Dynamics 365 Customer Service implementation eight weeks rather than twelve?
Six conditions, all of them properties of the project rather than things a partner can promise away. One support team on one channel, because email plus one live channel is a release and five channels at once is a programme. No migration beyond accounts, contacts, and open cases from a source that can be read directly, which removes the second load and the second reconciliation. No custom development, so the desk runs on configuration. The planning decisions already answered and signed, so week one is environments and observation rather than a discovery that ends in an argument. A single named decision maker reachable inside the working day, which is worth more calendar time than any other item on the list. And integrations limited to identity and at most one line of business system whose owner is genuinely on the project. If all six hold, the plan compresses to eight weeks because pieces of it disappear rather than because anyone works faster. If two hold, the honest answer is twelve.
What causes a Dynamics 365 implementation to overrun its timeline?
Implementations rarely overrun because the estimate was wrong. They overrun because something that was true in week one only became visible in week eight. The three root causes we test for on every takeover are requirements nobody owns, which surfaces in testing around week eight or nine and costs two to four weeks because every increment was built on the same wrong picture; technical debt in the environment being built into, which surfaces around week three or four the first time a change has to be promoted; and standard modules that were configured wrong or never configured and then built around, which surfaces around week four and is the cheapest of the three to fix while it is still a configuration change. Beyond those, the recurring delivery causes are data that nobody profiled before it was loaded, decision latency where a question asked on Tuesday is answered at next month's steering meeting, scope added mid build without displacing anything of equal size, the absence of a development, test, and production release path, and integrations to systems whose owners were never invited to the kickoff. On a healthy project the largest single risk is decision latency rather than engineering, which is why a partner in your own working day and language is worth calendar time rather than only convenience.
Our implementation is already late. Does the eight to twelve week plan still apply to us?
No, and neither does a fresh twelve week plan from anybody else. A project that has missed its date more than once has an undiagnosed cause, and any new estimate produced by the same process that produced the last one will miss as well. The first thing that has to happen is not a re-plan, it is a diagnosis: what is actively causing harm, what is going to fail next, and what must not be touched until the full picture exists. On our side that becomes a fixed scope reset phase rather than a new implementation timeline, with the Environment Audit Report and the Technical Debt Inventory at week two, a one page Reset Scope with a fixed price and a fixed end date, a working module in front of real users by week four, and a go or stop decision at week six that is yours to make. Where it is already an emergency rather than a slippage, environment lockdown and triage run in the first seventy two hours before any plan is written at all, because a timeline agreed on an environment that is still moving is a timeline that gets rewritten.
Why does our Dynamics 365 Customer Service routing not work?
In almost every case we inherit, one of three reasons. The attribute the route to queues decision list tests was never stamped, because work classification runs first and enriches the item, and routing that looks correct on screen never fires when the thing it reads is empty. Or there is no catch-all at the bottom of the decision list, so anything matching nothing drops into the workstream fallback queue and is forgotten. Or the queues were designed before anyone agreed who owns them, which is an organizational problem encoded as a routing problem. Before any of that, check the mailbox is approved and enabled for server side synchronization, because a mailbox that quietly failed test and enable stops email routing before a single rule is evaluated. Read the routing diagnostics on real traffic to separate an unmatched condition from a mailbox problem from an agent capacity problem.
How do you get agents to actually adopt Dynamics 365 Customer Service?
By treating adoption as a build decision rather than a training slot in the final week. Agents do not abandon a system for missing features, they abandon it because it costs them time on every case or because it drops them mid conversation. So we put agents in the room during configuration rather than only at user acceptance testing, design the case form for the first ninety seconds of a call, and settle the environment session and inactivity timeouts, the conditional access sign-in frequency, presence and capacity, and the browser policy on the agent fleet before go live rather than after the first week of reported disconnections. Single tab working becomes an explicit rule, since presence is per user rather than per tab. The pilot team gets real authority to change the build, hypercare runs for at least two weeks with a named person in the channel agents already use, and administrators are trained alongside agents so the configuration is not understood by one consultant who later leaves.
Do you work directly with our team or through our existing partner?
Both. Solzet works directly with in house teams to take over and finish a struggling implementation, and we also work on a B2B and white-label basis for other Microsoft partners who need to rescue a project without the client seeing a change of face. We are a Dynamics 365 Customer Engagement and Power Platform consultancy based in Yerevan, Armenia, delivering to clients and partners across Europe and the US, so nearshore and remote takeover work is how we normally operate.
Can you take over a Dynamics 365 migration that another consultancy stalled?
Yes, and it is a routine engagement for us, either directly or on a white-label basis for other Microsoft partners. The sequence is fixed. Anything still loading on a schedule is turned off first, because a migration that is still writing will undo corrections as fast as they are made. We preserve evidence at both ends, a copy of the target environment plus the staging tables and the job logs, and establish from the contract rather than the project plan the date the source system stops being readable, since that date sets the deadline for anything that has to be extracted again. Then we produce the reconciliation nobody produced: a per table comparison of target against source covering created dates, ownership, state and status, currency, and unresolved lookups. The mapping is rebuilt from what is really in both systems and the unmade business decisions are taken back to the people who can answer them. The load is rebuilt as something rerunnable from a service identity, rehearsed twice at full volume, and cut over against a timed runbook with a rollback point. We do not need the outgoing team to cooperate for any of this. Where the tenant is yours, tenant and Power Platform administration can be recovered through the tenant owner without the other side agreeing to anything, though one prepared hour with the people who built it is still the cheapest hour in the engagement.
How long does a stalled Dynamics 365 migration takeover take?
Six weeks to a rehearsed cutover on a typical Customer Engagement estate: days 1 to 3 to stop the writing and secure both ends, days 3 to 8 for the per table reconciliation, week 2 to recover the mapping and get the open decisions signed, weeks 2 to 3 to rebuild the load as something rerunnable, weeks 3 to 5 for two rehearsal loads at full volume, and weeks 5 to 6 for cutover and hypercare on the data. The first two weeks are fixed scope and end in a Migration State Report you keep whatever you decide next, so you are not committing to the whole six weeks to find out where you stand. It runs longer in two situations, both of which we would rather name at the start: where large tables have to be extracted again from a source with a closing read window, and where the business decisions that stalled the migration in the first place, on duplicates, ownership, and how much history is in scope, take longer to settle than any of the engineering does.
What is the difference between a stalled migration and a bad migration?
A stalled migration has not cut over. Both systems are still live, people are entering data twice, the cutover date has moved more than once, and the load either was never run at full volume or was run and never reconciled. A bad migration cut over and the data is wrong: reports do not match the system they came from, ageing looks impossible, duplicates multiply faster than anyone can merge them, and users have stopped trusting what the system tells them. The distinction matters because the deadlines are different. A stalled migration is racing the read window on the source system, so the first question is when that window closes. A bad migration is racing the environment's restore points and the point at which users have edited the migrated rows so heavily that reloading them would destroy real work. We run both, and the first week of each is spent working out which one you are actually in, because a team that treats a stalled migration as a bad one starts correcting data that is going to be loaded again anyway.
Implementation in crisis and need someone to take control?
Solzet takes over stalled and failed Dynamics 365 Customer Engagement and Power Platform projects, starting with a forensic audit and a stabilization plan and ending with a working, documented solution your team can own. Tell us what is broken and we will give you an honest diagnosis, directly or on a white-label basis for your team.