Writing

Productize this: building an AI product architect as a Claude skill

When I look at a product or a workflow, I usually start with the same questions. What is the person trying to accomplish? Where does the work get stuck? What could disappear if we designed it differently?

AI gives us more ways to answer those questions. It also makes it very easy to add complexity before deciding whether we need it.

I built ProductizeAI Architect to bring a more structured product conversation into that process. Give it a product description, a PRD or a workflow and say "Productize this." It proposes a redesign, explains the choices and identifies the first experiment to run.

The vision behind ProductizeAI Architect

I want it to help product managers, founders and product leaders decide what a product should become and why. The starting point is the user's outcome and the current friction. From there, it compares the simplest credible improvement with an AI redesign.

For each step, it asks whether the right mechanism is ordinary software, rules, predictive ML, generative AI or an agent. It then separates that technical choice from permission to act. A system might be able to take an action while still needing a human to approve it.

The output should make the before-and-after experience concrete. What does the user do differently? Which handoffs disappear? Who handles an exception? What happens if the model is wrong?

It also needs to account for the practical work around the model: integrations, review effort, maintenance and cost. A recommendation that saves someone time but creates a second job maintaining it needs to say so.

By default, the skill takes a quick pass through one workflow and up to three changes. A full blueprint goes deeper into architecture, autonomy, failure recovery, evaluation and economics. Either way, it should leave me with a bounded experiment and the question that could change the recommendation.

The product page is here.

How I turned that approach into a Claude skill

I started by making the method explicit. I wanted the same reasoning steps to carry across different briefs without having to explain my expectations from scratch each time.

The core instructions went into SKILL.md, with a name and description that explain when the skill should be used. I separated the more detailed decision rules and output blueprint into reference files. That keeps the main instructions readable while giving Claude a consistent framework to consult.

The decision rules cover choices such as rules versus AI, who owns a decision, what evidence is authoritative and how to evaluate a proposal. The blueprint sets out what a useful answer should contain: a product thesis, a before-and-after workflow, implementation options, dependencies, failures, costs and a first experiment.

Then I tested it on briefs with different constraints. I read the recommendations as a product manager would: could I understand the proposal, challenge its assumptions and work out what to do next?

When an answer missed something, I updated the reusable instructions. I wanted the next answer to improve too. These three tests were particularly useful.

Test 1: automating my renovation bookkeeping

My own receipt workflow was a good first test because the problem was concrete. I upload pictures to Google Drive, then later rename them, move them into folders and enter the company, description, date and price into my accounting spreadsheet. I often leave the second half for months.

The constraints were a minimal budget, low-code setup and low maintenance. The desired outcome was a spreadsheet I could send to my accountant.

The output proposed uploading once to a designated Drive folder. AI extracts the receipt fields. Automation adds a linked Google Sheets row and organizes the file. Uncertain values are flagged for review. It recommended a Make integration as the low-code route and a first experiment using twenty varied receipts.

I pushed back on the explanation. What exactly was a "receipt inbox"? Was this limited to renovations, or could it include medical receipts? What would it cost at an approximate volume? Which alternative replaced Make?

Those questions became improvements to the skill: define workflow components, clarify consequential scope, give useful cost estimates and compare credible implementation options with effort ratings. It also needs to explain what can be completed in ChatGPT and what requires external tools or development.

Test 2: redesigning a SaaS launch-readiness review

The second brief described a 120-person B2B SaaS company shipping about six features a quarter. A PM manages a 34-item checklist, chases owners on Slack and reviews open bugs without a consistent severity rubric. Legal review sometimes starts too late, help articles can describe an old UI, and launch conditions aren't followed up afterward.

The brief estimated six to ten hours of PM chasing per launch.

The output proposed a launch record backed by evidence from the systems that own the work. Its most useful decisions were about where AI belonged:

Workflow decisionRecommended approach
Legal reviewDeterministic triggers based on data-classification tags or schema changes; missing classifications escalate.
Checklist statusQuery Jira, CI and feature flags directly. Use AI to interpret ambiguous free-text replies.
Bug severityEstablish the rubric first. AI proposes severity with rationale; a human confirms it.
Help articlesUse AI to compare documentation and staging evidence and flag UI mismatches.
Launch decisionKeep go/no-go with the VP Product.
Post-launch conditionsTrack owners, due dates, evidence and escalation deterministically.

I checked the answer against a rubric and used the feedback to tighten the instructions. In particular, a missing severity rubric is a process gap to fix before asking a model to classify bugs. And an AI summary shouldn't replace a direct query to an authoritative system.

The proposed evaluation included total PM hours per launch, late Legal triggers, documentation-drift incidents, severity agreement and missed critical bugs. Those are measures for a pilot, not savings or reliability results I've already demonstrated.

Test 3: reviewing a product PRD

For the third test, I supplied the PRD for another project, Heard. This was a test of ProductizeAI Architect's ability to work through a product brief, with its requirements, constraints and trust expectations.

The output clarified the product's decision boundaries: which work needed model judgment, which claims needed verification in code and which decisions should stay with the product manager. It also recommended a smaller first pilot before expanding the collection capabilities.

That was useful because a PRD can make every requirement look equally ready to build. The redesign helped distinguish the core value from the dependencies that needed validation first.

I then chose to build the resulting skill. That implementation was a subsequent step; producing the architecture proposal didn't mean the integrations already existed or that the product was ready for production.

What the tests helped me improve

Across the three briefs, the same expectations kept coming up: make the recommendation concrete, name the assumptions, explain the trade-offs and define how we would know whether it worked.

The bookkeeping test exposed vague language and missing implementation detail. The launch test challenged the rules-versus-AI choices. The PRD test checked whether the skill could narrow a broader product into something worth validating first.

These were tests of the design recommendations. They don't establish production ROI. What they did give me was a way to improve the method and reuse it on the next problem.

That's what I want from ProductizeAI Architect: a product conversation that gets specific enough to change a decision, followed by an experiment small enough to run.

Try ProductizeAI Architect

The skill is public on GitHub: github.com/codingyogini/productizeai-architect.

Start with the workflow, who uses it, what frustrates them, the outcome you want and the constraints you have. Then ask: "Productize this."

I'd like to hear which recommendation you kept, which one you challenged and what you had to add before the answer became useful.

← All writing