Happy Saturday!
This week I continue to assess the differences between Claude and ChatGPT.
Last week I looked at how the two models compared in their analysis and critical thinking skills. I found that the two models performed similarly and wasn’t prepared to declare a winner or a loser. In fact, I felt that using both for the same task give me a wider perspective to consider.
This week, I put them through their paces by having them write a hypothetical white paper on reforming the tax code in Canada. On paper, this is where Claude should shine, as it tends to produce more natural-sounding writing.
However, ChatGPT has built-in web search capabilities which equips it to find recent sources.
Unlike last week, I have a clear winner for this experiment. Take a look:
The Prompts
For this experiment, I wanted to generate a white paper to explore alternatives to income taxes, with an eye to helping younger Canadians cope with housing and affordability costs.
Here are the sequence of prompts I used to get to an end product:
Prompt 1: Conduct baseline research
TOPIC=Alternatives to Income Taxes as a source of government revenue, INDUSTRY=Financial services, AUDIENCE=20-45 year old Canadians.
Use web search to identify 5-7 key challenges or pain points related to challenges high income tax rates pose for younger Canadians, especially with current affordability challenges. Summarize each in 1-2 sentences.Prompt 2: Find relevant trends
Research and list 3-5 current trends or innovations from other jurisdictions that have explored other forms of taxation. Include statistics or data points to support each trend.Prompt 3: Develop a compelling title
Develop a compelling title for the white paper that incorporates the need for tax reform in Canada and appeals to young Canadians. Create 3 options and briefly explain the rationale for each.Prompt 4: Draft executive summary
Once you pick your working title, move on to the executive summary:
Craft an executive summary (250-300 words) that outlines the white paper's main points, key findings, and value proposition for young Canadians.Prompt 5: Create a detailed outline
Create a detailed outline for the white paper, including: 1. Introduction 2. Background/Context 3. 4-6 main sections addressing key challenges and solutions 4. Case study or real-world example 5. Future outlook 6. Conclusion and recommendations. Provide a brief description of the content for each section.Prompt 6: Write the introduction
Write the introduction (500-750 words): 1. Hook the reader with a compelling statistic or scenario 2. Provide context for tax reform in Canada 3. Clearly state the white paper's purpose and what young audiences will gain 4. Include a brief overview of the main sections.Prompt 7: Build out the main sections
For each main section: 1. Start with a clear subheading 2. Present the challenge or issue 3. Provide in-depth analysis, including data and expert insights 4. Offer potential solutions or best practices 5. Include relevant graphics, charts, or diagrams to illustrate key points Aim for 1000-1500 words per main section.You will then need to prompt the model to build out each section, one at a time:
Please continue to write each section. You can write each section one at a time.Prompt 8: Create an infographic
Create a visually appealing infographic that summarizes the white paper's main points, key statistics, and recommendations.Prompt 9: Develop a reference list
Develop a reference list of at least 15 authoritative sources used in the white paper. Ensure proper citation throughout the document.Prompt 10: Write an author bio
Write an author bio (100-150 words) that establishes credibility and expertise on public policy.Prompt 11: Design a cover and table of contents
Design a visually appealing cover page and table of contents for the white paper.Prompt 12: Review, refine, and edit the draft
Review and edit the entire document for clarity, coherence, and consistency. Ensure it meets length requirements of 15 pages while maintaining high-quality, substantive content throughout.You will likely be presented with suggests for edits and an ask on whether you want the model to make those changes. I instructed the models to proceed with their recommended edits.
Prompt 13: Create a one-page summary sheet
Create a one-page summary sheet of the white paper, highlighting key takeaways and enticing young Canadians to read the full document.Prompt 14: Compile into one document
If needed, and depending on how you’re using the output from the previous steps, you may want to compile the work:
Compile all sections into a the cohesive white paper format you have produced. Please do not re-write the sections you have already re-written. Simply compile, following the table of contents have produced. Please do not re-write the sections you have already re-written. Simply compile, following the table of contents. Observations
So, how did they do?
Performance
I had to perpetually re-submit prompts with Claude, as it was constantly dealing with system performance issues. And that’s with me paying a pro license. ChatGPT had no problems keeping up with me. This became a source of deep frustration as I progressed through my prompts.
ChatGPT was the clear winner here.
Baseline research
Claude does not have web searching capabilities, relying on a knowledge base that is current as of April 2024. That didn’t stop it from conducting the baseline research required to write the paper:
ChatGPT pulled together a similar list, but added a couple more important points around limited investment opportunities, emigration of talent, delayed family formation, and the mental health strain.
ChatGPT wins out for its more comprehensive take and for providing me with sources I could verify.
Trends analysis
Claude surfaced three jurisdictional trends:
Estonia’s E-Residency Tax System
Singapore’s Goods and Services Tax
Norway’s Natural Resource Revenue Model
ChatGPT took a thematic approach to its response, clustering its analysis along five taxation models:
Flat Tax (Estonia)
Consumption-Based Taxes (New Zealand)
Environmental Taxes (Sweden)
Digital Service Taxes (France)
Wealth Taxes (Norway)
Again, ChatGPT comes out the winner for offering range and sources.
Executive Summary
This is where Claude starts to shine. I find its executive summary easy to read and compelling all the same:
By contrast, while there’s nothing wrong with ChatGPT’s executive summary, it just reads like a string of patterns — predictable. It has that “AI feel.”
Ultimately picking a winner/loser here is a matter of preference. My personal preference is Claude, for having a slightly more natural feel.
Introduction
Both models opened their introductions with a story of a fictional Canadian persona. Both follow the instruction of the prompt well, providing context for tax reform, stating the white paper’s purpose, and providing a preview of the main sections. On a qualitative note, ChatGPT’s version reads like a first-year university student trying very hard to sound smart. Claude’s version has an understated confidence that feels more natural and professional, to my eye. You can read both of their introductions in the final output files.
Main Sections
Claude was much more faithful to the directions of the prompt when writing each section. Claude also offered sample charts and graphic solutions with proposed Mermaid code to build those charts.
Here is a sample graphic idea it offered when illustrating how the current tax structure affects housing affordability. Nothing ground-breaking, but I appreciated the effort and the seeds of a concept I could improve if I wanted.
It also offered me key stats (which I would need to verify, obviously), and comparative tables.
That said, because of system performance overload, I was limited to succinct answers, which really hampered the end product.
For each section, I got headers and bullets. Take a look. This had a negative impact on readability. Each section was tough to follow. It wouldn’t provide me with contact for each heading and supporting bullets. I should note that in my early testing when the system was performing normally, I did get quality written-content, but I couldn’t replicate this despite multiple attempts at different times of the day. If I was working on a deadline, I’d be in rough shape.
So, this one is tough. If I can catch Claude on a good day, I might use it. But ChatGPT wins this round for at least showing up reliably. Besides, it did offer ideas for illustrations and charts. I also like that it included expert quotes from actual prose to set up each section.
Case Study
ChatGPT picked Sweden’s carbon tax for its case study. An interesting choice, given how much of a political hot potato the carbon tax has been in Canada, with many progressive leaders doing a 180 on it. That said, I wasn’t expecting ChatGPT to give me sage political advice. It otherwise gave me what I was looking for, outlining the challenges Sweden faces, the policy design, the outcomes the policy achieved, challenges it faced and lessons for Canada.
Claude picked Estonia for its case study. The output was garbage. I asked for a case study and the only context I got was this:
Estonia's transformation from a post-Soviet state to a digital leader offers crucial lessons for Canadian tax reform. Their e-Tax system processes 98% of tax returns in under 3 minutes, with 99% of government services available online.
Followed by a bunch of Mermaid code, tables and bullet points. Here were the lessons for Canada:
You can see the output of here.
Pretty underwhelming. ChatGPT is the clear winner here.
Infographic
Ouch. This one is rough. ChatGPT’s attempt is simply not useable:
“Introduce income taxes for young tax credits.” WTF.
Claude’s attempt was better, but far from useable.
A big fail for both models here.
Sources
You can find the full list of sources Claude provided me here. The problem? Many of these simply don’t exist. When I asked Claude to provide me with links to validate these sources, it fessed up to having fabricated them.
When I asked ChatGPT to provide me with specific sources, it did. I was able to click out to each link.
ChatGPT is the clear winner here.
Cover Page
This one has trade offs. ChatGPT got creative with its attempt at a visually compelling design. Claude was far from creative, but it was at least legible. I’ll give Claude the win here.


The Takeaway
I felt like I got two entirely different versions of Claude. When I first designed and tested my prompts, the output I got from Claude was wonderful.
Take a look at how it originally wrote one of the main sections. It identified key challenges that read very well, with just enough detail. However, I could not replicate this despite many attempts.
Only 24 hours later, when I re-ran my prompts, I got the results you see above. It looked promising, producing an executive summary and introduction on par with what I had originally experienced. And then, it just fell off a cliff.
Other than system performance, the only main difference I can spot is that in my original test, I had an extremely long prompt that I ultimately decided to break up into the 14 prompts you see above. Perhaps the mega-prompt helped prime Claude much better than the sequence.
By contrast, ChatGPT reliably gave me consistent results every time I tested it.









