A skill became an artifact that the pipeline generates, archives and consults on its own. This is automation with supervision at the right points.
Image generated with Gemini
In Part 1, I wrote about building skills so the AI understands your project. In Part 2, the harder problem: writing a skill so the AI understands you.
Part 3 is different. It picks up where Part 2 left off, with the infrastructure needed to catch up with the partnership, and it ends at a point that exceeded my own expectations: the pipeline learning to fix its own errors and document what it learned.
August: The Month of Volume
Source Genesys (SG) is a technology that creates vertical SaaS platforms. It’s fully automated with Cursor and we have already several platforms in production. SG has a CLI command that automates the process of programing with a custom pipeline CI/CD. August was the time to test and design the skills for Cursor and Claude.AI. There was a lot of documentation to analyze the quality and security of the code.
The August report logged 507 executions across 833 sessions, with a success rate of 85.4%. That is 10 active platforms, 10 generators running in parallel. All the data comes from my own work sessions.
The number that matters most isn’t the total. It is in the distribution. AAA (the main tenant system controller) led in publications with 47, followed by MDS with 39. Together, the two platforms accounted for more than half of the month’s total volume. That is no accident; it reflects where development was concentrated.
The peak activity hour is 2 p.m., with 72 executions. But real behavior shows up when you look at the full curve: there is significant activity at 10 p.m., midnight, and 1 a.m. Work does not stop when the day ends. The pipeline is available for as long as you are awake, and that changes how you distribute your attention throughout the day.
How Activity Spreads Across the Day
The chart below shows SG CLI automated command executions per hour in August. The 2 p.m. peak (72 executions) coexists with consistent late-night activity, making it clear that the pipeline operates as an extension of the work rhythm, not a replacement for business hours.
Daily Activity by Hour
The Price of Volume: Failures and What They Reveal
High volume brings failures. 74 executions ended in error during the month. The overall rate landed at 85.4%, technically within the limit, but what matters is the daily variation, not the average. The daily success rate chart swings between 100% and 25% throughout the month, with perfect peaks followed by rough days. This is not progressive degradation. Episodic instability is harder to diagnose than steady decline.
The failure heatmap by platform and command reveals the pattern clearly. For ActivateObservability, a command to self-publish observability on Grafana Cloud without human intervention, MDS holds the highest absolute number of failures: 12 occurrences. That is because it is the base system and template for all the other platforms. The number does not come from a difficult platform; it comes from a command with inconsistent behavior. The failure happened during activation, not in the platform itself. When the same command fails in [PIX]BOL, CRMW, AAA and Entity, the hypothesis of an isolated issue becomes unsustainable.
ActivateObservability ended August with a 48.6% failure rate. TestAPI came in at 66.7%. Together, those two numbers point to where the pipeline still has work ahead: observability activation and API testing, which are customized per platform, are the highest-pressure points in the system.
Failure Rate by Platform: Where the Pipeline Still Has Work to Do
The chart below shows the failure rate of my project by platform in August. MDS leads with 25.7%, the highest value in the ecosystem, and it is the base template that feeds all the other platforms.
Fail (%) by Platform
MDS’s 25.7% failure rate demands attention for two reasons. The first is volume: 39 publications in the month means each failure has a real-time cost. The second is the MDS platform’s template nature. Problems here can propagate to every generated platform.
But the number doesn’t show the full picture. MDS is also the platform where development is most intense, where new patterns are tested first, and where the pipeline receives the most significant changes before they are propagated. A high failure rate on a platform under active development is not the same as one on a stable platform.
What Skills Changed in the Pipeline
Part 1 of this series described how skills were built by hand: one Markdown file per domain, loaded in the right context, teaching the AI what it could not learn from generic training. In August, that architecture evolved in two directions.
The first direction was scale. With more than 500 active skills covering everything from system identity to password patterns, CI/CD, observability and Cursor behavior, the volume of context available per session grew. That directly affects the quality of generated prompts: when the AI knows the canonical pattern, the prescriptive prompt gets shorter. Less room for Cursor to improvise.
The second direction was unexpected.
Auto Fix: When the Pipeline Learns to Write Its Own Skills
On August 15, the pipeline log showed something different in AutoFix, an AI module using Cursor and several skills. AutoFix had run, corrected a TypeScript error in AAA, and automatically generated the skill 31-sg-ts2307-module-not-found.mdc. No manual intervention. No additional prompt.
The name that came out of it was SG AUTO FIX.
The mechanism worked like this: when AutoFix resolved an error with no prior record, it generated a skill with the problem pattern, the affected platform context and the solution applied. The skill was saved with a 15-day TTL, a sequential number, and made automatically available for future sessions.
Then the pipeline gained a distinction that, in practice, is the most important of all: before applying any fix, AutoFix had to classify the file with the error. If the file was generated by SG and had a corresponding template mapped in SG code, the fix had to go into the template, not the generated file. Fixing the output without fixing the source would mean that the next time the generator ran, the error would come back.
That changed the nature of AutoFix, from remediation to source-level correction.
To make sure no template fix went through without review, SG AUTO FIX was configured to fire a notification whenever it changed an SG template, with the full diff of what it was and what it became. The learning loop gained a mandatory point of human supervision at exactly the spot where the risk is highest.
What High Volume Revealed About the Pipeline
With the high volumes of August’s data consolidated, the dynamic had shifted: Cursor was generating its own fix instructions, and Claude had become the second opinion, the one I’d check before approving what the pipeline had decided on its own.
PublishDirect (the CI/CD that publishes artifacts fully automated) accumulated 924 minutes in the month, more than 15 hours of build time, over 10 times the next most expensive command. It consumes the most time because it does the most work. The point is not to reduce that number, but to understand what is inside it: every minute of a successful PublishDirect is a platform running without manual intervention.
Average build time per platform ranges from 0.37 minutes for ADV to 4.2 minutes for CRMR. That 11x variation between the extremes is not necessarily a problem; more complex platforms have longer builds. What matters is the P90: CRMR and MDS have the highest P90 in the ecosystem, which means a significant fraction of builds on those platforms is much slower than the median. When P90 is much higher than P50, the pipeline is unstable, not slow.
Put it this way: if P50 = 8 min and P90 = 22 min, there is a 14-minute gap.
That means every time something leaves the happy path, the build takes nearly 3x longer. The pipeline is not slow; it is unpredictable.
A slow pipeline would have P50 and P90 both high.
An unstable pipeline has P90 much higher than P50; in this case, CRMR and MDS.
The build minutes consumed per day chart shows that July 31 burned nearly 160 minutes, followed by a sharp drop. Cross-referencing with session volume, it was the highest-activity day of the period. There were three rounds of complete rebuilds across multiple platforms, in sequence, to investigate and consolidate template errors. Each round produced a block of logs that AI automatically analyzed, generated a targeted prompt and applied the fix, with manual, interactive work alongside Cursor, case by case.
What Is Still Human Work
With an automated pipeline, active observability, AutoFix generating skills and template notifications, the obvious question is: what still requires a human in the loop?
The honest answer is more than it seems, and at exactly the most important points.
A green build is not validation. AutoFix applies, compiles and reports. End-to-end testing on each platform is still done manually, and that step has not been automated. A build that passes with broken code is worse than a build that fails with correct code, because the failure is visible and the bug is not.
Template fixes require review. When SG AUTO FIX changes a template file, the notification is sent with full diff. Accepting that change without reviewing it would mean delegating to the agent a decision that affects every generated platform.
Conclusion
When I wrote Part 1, a skill was context. In Part 2, it was posture calibration. In August, it became an artifact that the pipeline generates, archives and consults on its own.
This is not full automation. It is automation with supervision at the right points: template fixes, validation testing and prompt scope. The AI executes more and decides less; the developer reviews less routine work and makes more decisions.
The number that sums up August is 85.4% success across 507 executions. But what that number represents is different from what it would have represented six months ago. In January, 85% success would have meant 85% of a small task completed.
In August, it means 85% of 10 platforms running with no manual deployment and no manual observability setup, and, for the first time, a pipeline that knows how to document its own fixes.
Automate with AI. But do not give up reviewing the points that matter. That part is still your job.
Automate with AI. But do not give up reviewing the points that matter. That part is still your job.



![How SafeStyle Shipped 100,000+ Orders in Under a Year with ShipBob [Case Study]](https://www.shipbob.com/wp-content/uploads/2026/10/ba05c72624a0a82d7de08863e1e0a1a2.jpg)






