Claude + ICM at home for Copilot at Work
TLDR: As is the case for several of us here, we’re only allowed to use Microsoft Copilot at work. I wanted to take an ICM approach to a large scale contract audit and validation assessment using Copilot, so designed the workflow at home with Claude to generate a “work pack” of agents, Notebook instructions, and prompts to install at work.
One key learning for me was that even though ICM shaped it well, I hit multiple session limits and an embarrassingly large number of tokens polishing the specification in Claude before implementing anything in Copilot.
The problem:
Part of my day job includes keeping a large ICT contracts register honest against the actual evidence sitting in SharePoint: renewal dates, pricing, which document supersedes which, etc etc. Often a case of reporting vs reality drift, plus the register has multiple scenarios difficult for a direct approach to analyse e.g. vendors that are intentionally duplicated as separate rows in the register (single reseller, multiple individual/separate software contracts), large single contracts that cover multiple major agreements, and invoices, managed separately, that extend a contract term based on annual payments.
I don’t have a team supporting me, I’m ‘it’ and contract management is just one part of my job so I’m having to juggle over a hundred vendors along with managing finances and other commercial requirements for a large IT department.
At work the rule is that it's Microsoft Copilot or nothing, so the design has to survive being executed manually inside Copilot Chat, Notebook/Researcher/Analyst, and Agent Builder, with a documented fallback for anything not currently available within our tenant e.g. we don’t have access to Cowork (yet).
In my initial attempts with Copilot, using the “WorkIQ” functionality enabling it to reference SharePoint and internal documentation (including emails and Microsoft Teams), it pulled in old workbooks, unrelated vendor documents, and just a whole bunch of wrong data – despite trying to lock it down – and royally messed things up.
And, of course, none of the actual data can come home, it can’t be on my home computer so that I can build a solution from what I’m actually working with. To get around that, I got Copilot to analyse the problem from the data I provided it with and produce a document based on the multiple contract scenarios needing to be addressed. I then used that to provide Claude with the problem statement for the development of the project plan and ICM structure.
The approach:
The actual register is maintained in Monday.com as it syncs with other department registers into a master record that the Legal team manages. I therefore have to export the register to Excel and work from there. I use Copilot to assist with Power Query creation in Excel, but the process Claude has come up with covers the use of a mix of tools:
  • Excel/Power Query for managing the base data.
  • Two Copilot ‘declarative’ agents return findings from the SharePoint vendor folders (100+) only for rows that exist in the spreadsheet (the register only tracks contracts over a set value, we have a lot below that which aren’t reported at the same level so are excluded from the process).
  • Everything is fail-closed, missing or conflicting evidence gets flagged, with some serious rules in place about not guessing or assuming anything, or going walkabout to find some random document that seems to fit the requirement and building it in (true story 😐). Every confirmed value has to trace back to a specific document and page/locator.
  • Complexity is tiered (simple / standard / complex) so the ‘expensive’ tools (Copilot Researcher and Analyst) are only used for the contracts that actually need them.
That last point was a real ‘gotcha’ for me as until I got Claude to determine the best approach via the available Copilot tools, I hadn’t realised that Researcher and Analyst have a “combined” monthly limit of 25 queries! Monthly! And people complain about Claude limits.
Copilot’s declarative agent instructions also cap out at 8,000 characters per field (and the sum of fields can still trip errors under that – Claude designed a range of tests to determine how/whether we would have issues here and then built around it), and whilst Copilot Notebook states up to a maximum of 300 references, there are a bunch of caveats around that. So instead of linking to SharePoint folders which can contain 10s of documents per vendor (much more if invoices drive renewals in addition to ad-hoc/monthly/non-renewal based charges), anything needing to be considered as part of the review gets added via AI assessment as an individual reference, not left to a folder scan.
The structure:
The whole thing is built as an ICM workspace (of course 😊), numbered stages, each with its own folder and an output folder within each stage where the actual deliverable lives, gated by review before the next stage starts. The folder structure carries the architecture as per process. As I’m producing everything at home to take to work, I’ve called it the ‘Contracts Validation Factory’:
ContractsValidationFactory/
shared/ - the standing design brief and references that every stage designs against
stages/
01_data_model/ - data dictionary, workbook spec
02_controls/ - evidence rules, tiering, quota budget
03_agents/ - the declarative agents and instruction sets
04_prompts/ - prompt catalogue per tier, output-format contract
05_test_pack/ - test cases to check against limits, correct data returns etc
06_release/ - the work-side implementation tasks and operating procedure
The process:
All building gets done at home in Claude which produces the ‘work pack’, a Claude output folder that carries only what's needed to run the process at work:
· a day-one capability checklist,
· the run procedure (in sequenced markdown files),
· paste-ready prompts and agent text,
· a handful of graded test cases built around the register's trap patterns,
· an Excel workbook created by Claude into which I feed the contract register data, containing a bunch of formulas and extra fields to capture the output from the agents etc,
· and, as the only file that I return home, a test-results markdown file that I ‘fill in the blanks’ for Claude based on the results at work – no contract data involved, all process and outcome related – for Claude to analyse and fix or build on.
If I run into issues with the process at work I don’t try to fix it unless it’s something simple that won’t have downstream impact, I log it in the ‘test results’ file and move on to the next step, Claude and I work it out when I get the test results home.
Where it went sideways, briefly (sort of):
In one of the review passes, Claude confessed to introducing several defects. I subsequently got caught in a loop of Claude reviews generating output that ChatGPT analysed and produced its own analyses that Claude then reviewed, accepted that ChatGPT had identified more issues, and produced a further report for subsequent review … and back and forth we went. The project became more about proving itself correct without actually testing anything against the contract environment. It was an embarrassing sidetrack consisting of multiple review passes of the same specifications (Claude may or may not have used the word ‘novel’ in terms of amount of wasted output!). Lesson learned: Claude will happily keep making the design ‘more consistent’ forever. Determine “done” for the initial build, test it against the actual data, iterate through the process rather than trying to get things perfect in the home environment without the actual data to test against.
Current state
I’ve commenced implementation at work and in setting up the working spreadsheet that Claude built, identified a couple of process issues that hadn’t been captured in the initial analysis. I’ve got the updated work pack ready to go, just needing a gap between meetings and “urgent” stuff to implement and test.
8
2 comments
Mike Burford
3
Claude + ICM at home for Copilot at Work
Clief Notes
skool.com/cliefnotes
What we give away free beats most paid courses. Build durable AI systems with a Marine vet and Edinburgh researcher. 40+ lessons, growing.
Leaderboard (30-day)
Powered by