How I Research Company Reviews — Full Methodology
Quick Answer: Every company report I publish goes through a structured, multi-step process: up to eight specialized AI research agents, fifteen-plus public databases, review analysis with independent dual-classifier verification, dual AI cross-checking, a 43-point quality gate, and a liability review. This page explains the full process — how the research works, what data sources feed each section, how I separate real program reviews from sales-call reviews, and what the limitations are.
I publish AI-assisted research reports on companies in the debt relief, credit repair, and consumer finance space. These reports compile publicly available data — corporate filings, court records, government complaints, consumer reviews, and website compliance — so consumers can make informed decisions. Here is exactly how every report gets built, from start to finish.
Why I Built This Process
When consumers search “Is [Company] legit?” they find a mix of affiliate review sites (paid to send leads), AI-generated content (unsourced assertions echoing each other), and genuine reviews that may describe a phone call rather than actual program results.
The Problem: A consumer researching a debt settlement company might read a “review” site that was paid by that company, see 4.7 stars from thousands of reviews that mostly describe a phone call, and find AI-generated blog posts making corporate ownership claims with zero state filing evidence. None of that helps them make an informed decision.
I built this research process to produce something different: a structured, sourced, verifiable compilation of what the public record actually shows about a company — the good and the bad, weighted equally.
The Research Team: Eight Specialized AI Agents
Each report starts with up to eight specialized AI research agents working in parallel. Each agent has a defined scope and a specific set of databases to search. They do not communicate with each other — they report independently to avoid groupthink or confirmation bias.
Agent 1: Corporate Records Investigator
Traces the company’s legal history through state corporate filings. This agent follows chains of evidence: find the legal entity, find the principal’s name, search that name for other entities, search those entities for shared addresses. What emerges is a corporate genealogy — who owns what, who is connected to whom, and how the company has changed over time.
Databases searched: Florida Sunbiz, state secretary of state databases (all 50 states where relevant), OpenCorporates, SEC EDGAR (for publicly traded companies), FDIC BankFind (for banking partners), and NMLS Consumer Access (for lending licenses).
Critical rule: Every corporate connection claim must be verified through state filings — not third-party blogs. If a review site says “Company A is owned by Company B,” this agent checks whether that claim is supported by actual corporate records. Claims that cannot be verified are flagged as unconfirmed.
State Registration Verification: Beyond the initial corporate research, a separate step deploys parallel search agents across 15 or more state secretary of state databases to verify where the company is actually registered to do business. Each state is searched individually and the result recorded as Found, No Results, or Unable to Search (for states with CAPTCHA barriers or access limitations). Reports document the exact number of states searched versus states that could not be accessed — no assumptions are made about states that were not searched.
Agent 2: Court and Enforcement Analyst
Searches federal court records and government enforcement databases for lawsuits, consent orders, and regulatory actions involving the company and its principals.
Databases searched: CourtListener (federal court records via authenticated API), FTC enforcement database, CFPB enforcement actions, and state attorney general press releases.
This agent also searches principal names individually — an executive’s involvement in a previous enforcement action at a different company is relevant context for consumers.
Agent 3: Consumer Review Analyst
Collects and analyzes consumer reviews across all major platforms. Agent 3 does not just read the most recent reviews — it pages through multiple pages on each platform to collect reviews spanning different time periods, typically 30 to 50 reviews per platform. This avoids recency bias and captures both older and recent experiences.
Platforms searched: Trustpilot, Better Business Bureau, Google Reviews, BestCompany, and PissedConsumer.
Agent 3 classifies every collected review into one of three categories: Call/Intake, Program Performance, or Ambiguous. It also writes the raw review text (without its classifications) to a separate file for Agent 8 to classify independently. The full classification methodology is described below.
Agent 4: Web and Background Investigator
Researches leadership backgrounds, marketing practices, and third-party claims about the company. This agent is specifically tasked with finding claims made about the company on other websites — and checking whether those claims are accurate.
Databases and tools searched: Facebook Ad Library (active ads and third-party advertisers), Wayback Machine CDX API (website history and changes over time), Glassdoor/Indeed (employee reviews), and general web searches for news coverage and third-party claims.
Why the Facebook Ad Library matters: It reveals which companies are actively advertising and whether third parties are running ads using the company’s name. This can uncover lead generation networks, celebrity endorsement campaigns, and multi-brand operations that consumers would not otherwise discover.
Agent 5: Website Compliance Analyst
Reads the company’s entire website with a regulatory compliance lens. For debt settlement companies, this means checking for all eight FTC Telemarketing Sales Rule (TSR) required disclosures. For other company types, the checklist adapts — FDCPA for debt collectors, TILA for lenders, CROA for credit repair companies.
This agent documents what the company discloses and what it does not. Disclosures that are present get credit. Disclosures that are absent are noted as gaps — not as accusations.
Agent 6: Complaint and Response Analyst
Analyzes formal complaint data from the CFPB Consumer Complaint Database and the Better Business Bureau. But the complaint count is only half the picture — this agent also reads how the company responds to complaints.
Every sampled company response is classified by tone: Resolution-Oriented (reports use “Tried to Fix It”), Procedural/Template (“Generic / Copy-Paste Reply”), Defensive (“Defensive / Blamed Customer”), or Non-Responsive (“No Response”). A company that consistently offers specific resolutions to unhappy customers tells a very different story than one that sends template responses.
Agent 7: Nonprofit Financial Analyst (Conditional)
This agent deploys only when the company under review is a 501(c)(3) nonprofit organization. It retrieves and analyzes the organization’s IRS Form 990 tax filings via the ProPublica Nonprofit Explorer API, extracting executive compensation, program service revenue, total assets, and how the organization allocates funds between program services and administrative overhead.
For-profit companies skip this agent entirely. When it does deploy, its findings feed into dedicated report sections on financial transparency and executive compensation that are unique to nonprofit reviews.
Agent 8: Independent Review Classifier
This agent exists for one purpose: to independently verify Agent 3’s review classifications. It receives only the raw review text — no access to Agent 3’s decisions — and classifies each review using the same fixed criteria. The two sets of classifications are then compared to measure agreement.
This is the inter-rater reliability step described in detail below.
Data Sources
Every report draws from publicly available databases. No private records, leaked data, or insider information is used.
Government and Legal Records
- CFPB Consumer Complaint Database (API)
- FTC enforcement actions
- State attorney general press releases
- CourtListener — federal court records (authenticated API)
- SEC EDGAR — corporate filings for publicly traded companies
Corporate and Business Records
- State secretary of state databases (all 50 states)
- OpenCorporates — global corporate registry
- NMLS Consumer Access — lending licenses
- FDIC BankFind — bank verification
- ProPublica Nonprofit Explorer — 990 tax filings
Consumer Review Platforms
- Better Business Bureau (ratings, complaints, responses)
- Trustpilot
- Google Reviews
- BestCompany
Web and Marketing Intelligence
- Facebook Ad Library (active advertisements)
- Wayback Machine (website history via CDX API)
- Company websites (compliance analysis)
- Glassdoor / Indeed (employee reviews)
Connection Analysis: How I Trace Corporate Relationships
Many companies in the debt relief space share principals, addresses, or corporate histories with other entities. A consumer signing up with “Company A” may not realize it was created by the same people who ran “Company B” — which may have enforcement actions against it.
I use a tiered evidence framework to classify every connection claim:
| Tier | Evidence Level | What It Means |
|---|---|---|
| Tier 1 | Confirmed | Court “fka” language, Articles of Amendment, or same entity filing ID. The connection is a documented legal fact. |
| Tier 2 | Strongly Linked | Same registered agent + same principal + same address. Or entity dissolved within 90 days of new entity filing. Multiple independent indicators point to the same conclusion. |
| Tier 3 | Associated | Same address but different principals, or shared phone/email domains. A weaker signal that warrants mention but not a definitive connection. |
What Is NOT a Connection: Name similarity alone. Two companies sharing a common word (like “Century” or “National”) are not connected unless they share principals, addresses, or filing history. This is the single most common error in third-party “reviews” of debt relief companies — and the rule I enforce most strictly.
Every connection in every report is labeled with its evidence tier. Readers can see exactly how strong the evidence is and judge for themselves.
Review Analysis: Separating Calls From Results
Why Review Ratings Can Be Misleading
Most debt relief companies have thousands of online reviews with high ratings — 4.5 stars or above on Trustpilot, Google, and similar platforms. Those numbers look great until you read the actual reviews.
Debt settlement programs typically last 24 to 48 months. During that time, consumers pay fees, their credit scores drop, and creditors may sue. But many companies send review requests within days of enrollment — at the moment of peak hope, before any of those consequences arrive.
The result: a company’s published rating might be 4.7 stars, but if 75% of those reviews describe the phone call and only 25% describe the actual program, consumers are making decisions based on incomplete data.
My Classification System
Every review in my sample gets classified into one of three categories using fixed criteria defined before the data is collected:
(settlements, fees, timeline, credit impact)
Most informative for consumers.
(named agent, phone call, no outcomes)
Call/Intake Review (requires 2+ indicators)
- Names a specific agent or consultant
- Describes a phone call, consultation, or initial meeting
- Mentions being “walked through options” or “feeling understood”
- Contains no reference to program outcomes, fees charged, or timeline results
- Posted within 30 days of self-reported enrollment
Program Performance Review (requires 1+ indicator)
- Mentions specific financial outcomes (debt settled, amount saved, fees paid)
- References program timeline or duration experienced
- Describes credit score impact during or after the program
- Reports settlement-specific events (creditor lawsuits, 1099-C tax forms)
- Discusses cancellation experience or post-cancellation events
Ambiguous reviews — those that do not clearly fit either category — are classified separately and excluded from the adjusted rating calculation. They are still counted and reported.
Precedence Rule: If a review contains both call/intake indicators AND program performance indicators, it is classified as Program Performance. The rationale: a review that mentions both the phone call and the settlement outcome contains outcome data, which is the more informative signal for consumers.
How Reviews Are Collected
Reviews are not cherry-picked. For each platform where a company has reviews, Agent 3 reads through multiple pages to collect 30 to 50 reviews spanning different time periods — not just the first page of most recent reviews. If a platform has fewer than 50 reviews total, all reviews are collected.
Why Not Larger Samples?
The statistically ideal sample for a company with 16,500 Trustpilot reviews would be 376 reviews (Cochran’s formula: 95% confidence level, ±5% margin of error). That math is straightforward. The review platforms, however, impose access restrictions.
Platform Access Realities: Trustpilot restricts unauthenticated browsing beyond approximately 10 pages (~200 reviews) with a login wall. Other platforms use pagination limits, dynamic content loading, or anti-bot measures that prevent collection at scale. These restrictions are not unique to this research process — they affect anyone who tries to read reviews beyond the first few pages.
Rather than claim sample sizes that cannot be reliably collected, this process is transparent about what review platforms allow. The reviews that are accessible are collected and analyzed. The number collected, the total available, and any access limitations are documented in every report.
What This Means for Precision
Margin of Error: With 30-50 reviews per platform, the margin of error for percentage estimates is wider than it would be with 376 — roughly ±12-14% instead of ±5%. Every adjusted rating includes a 95% confidence interval that reflects the actual sample size collected. A wider interval does not mean the analysis is wrong — it means it is less precise. The direction of the finding — whether program-outcome reviews rate significantly differently from call reviews — remains meaningful and reportable.
Every report discloses: the number of reviews collected per platform, the total reviews available, any access limitations encountered, and the resulting confidence intervals. Readers can evaluate the statistical strength for themselves.
The dual-classifier approach compensates. While a smaller sample means wider margins on the rating estimate, the inter-rater reliability measurement (Cohen’s kappa) is independent of sample size — it measures whether the classification criteria are being applied consistently. A kappa of 0.75 on 40 reviews is just as meaningful as a kappa of 0.75 on 376 reviews. The classification system’s objectivity is verified regardless of how many reviews are collected.
Reference: Statistical Sample Size Formula
Base sample: n = (Z² × p × (1-p)) / E² = 385 reviews
This assumes: 95% confidence level (Z = 1.96), maximum variability (p = 0.5), and ±5% margin of error (E = 0.05).
Adjusted for population size: nadjusted = (385 × N) / (385 + N – 1)
Where N = total reviews on the platform.
This formula defines the target sample size. As automated sampling tools are developed, sample sizes will scale toward these targets. Current reports document their actual sample sizes and resulting precision.
Independent Verification
Dual AI Review Classification
Classification is subjective. Two reviewers reading the same text might disagree about whether it describes a phone call or a program outcome. To measure and account for this, I use two independent classifiers.
How It Works: Agent 3 (the Review Analyst) collects and classifies the review sample. Agent 8 (the Independent Classifier) receives only the raw review text — no access to Agent 3’s decisions — and independently classifies each review using the same criteria. Then I compare their classifications.
Cohen’s Kappa
I measure agreement between the two classifiers using Cohen’s kappa (κ), a statistical measure that accounts for agreement that would happen by chance:
κ = (Po – Pe) / (1 – Pe)
Po = observed agreement (how often they actually agreed)
Pe = expected agreement (how often they would agree by chance alone)
| Kappa Range | Agreement Level | What It Means |
|---|---|---|
| 0.81 – 1.00 | Almost Perfect | Classifications are highly consistent — high confidence in the adjusted rating |
| 0.61 – 0.80 | Substantial | Good agreement — adjusted rating is reliable |
| 0.41 – 0.60 | Moderate | Meaningful agreement, but many reviews are genuinely hard to categorize |
| 0.21 – 0.40 | Fair | Low agreement — the boundary between “call review” and “program review” is blurry for this company |
| < 0.21 | Poor | Classifiers disagree frequently — adjusted rating should be interpreted with significant caution |
Kappa is computed for every report. The full kappa score and contingency matrix are available in the research files. Reports themselves use plain language — “two AI reviewers independently sorted every review” — with a link back to this page for readers who want the statistical details. Low kappa is not a failure — it is data. It means the company’s reviews genuinely mix call and program content, which is itself useful information for consumers.
Resolving Disagreements
When the two classifiers disagree on a review:
- One says ambiguous, the other commits to a category: I use the non-ambiguous classification (the classifier who committed had a reason)
- One says call/intake, the other says program performance: The review is classified as ambiguous (genuine disagreement — I do not force a category)
- Both agree: That classification stands
The Adjusted Rating
After classification and disagreement resolution, I compute two separate ratings for each platform:
Program Performance Rating (called “Results-Only Rating” in reports) = average star rating of reviews classified as describing actual program outcomes.
Call/Intake Rating (described as “reviews about the sales call” in reports) = average star rating of reviews classified as describing the initial call experience.
The gap between the published rating and the Results-Only Rating tells consumers how much of the company’s rating comes from program results versus sales call experiences. Reports present these numbers in plain language; the technical methodology is documented here.
Confidence Intervals
Every adjusted rating includes a 95% confidence interval:
CI = rating ± 1.96 × (s / √n)
Where s = standard deviation of program review star ratings, and n = count of program reviews.
If only 8 program reviews exist out of 40-50 collected, the confidence interval will be wide (possibly ±1+ stars). I report that honestly. A wide CI means: “There are not enough program-outcome reviews to be precise about this company’s actual service quality.” That is itself a finding — and it often reveals that the vast majority of a company’s reviews describe the sales call, not the program.
Dual AI Cross-Check (GPT Review)
After the report is written and fact-checked, the entire document is submitted to a second AI engine (GPT, by OpenAI) for independent review. GPT checks for:
- Factual accuracy — do numbers and dates match known public records?
- Logical consistency — do sections contradict each other?
- Missing public data — is there significant information about the company not included?
- Fairness — is any section framed in a way that implies wrongdoing rather than presenting data neutrally?
- Connection verification — are corporate ownership claims backed by state filings or only by blog posts?
- Entity completeness — is every named person and company properly tagged?
GPT’s findings are triaged into three categories: Fix (real errors — must be resolved), Improve (strengthens defensibility — usually implemented), and Skip (stylistic preferences — no action needed). Every Fix item must be resolved before the report can proceed. The GPT review artifact is saved and referenced in the quality gate.
Quality Control: The 43-Point Gate
No report reaches the site until it passes every check in a 43-point quality gate. If any check fails, the issue is fixed and the entire gate runs again from the top — because a fix in one area can introduce a new issue elsewhere.
The gate is a blocking loop. There is no override, no exception, and no “close enough.” Every check must show PASS before the report is uploaded.
The gate runs in three phases: an automated script handles 33 mechanical and quality checks in seconds (HTML structure, source link validation, CFPB data consistency, star rating arithmetic, connection tier labels, and more). Then seven checks requiring LLM judgment are evaluated manually (source verification, corporate connection evidence, entity type profiling, state registration confirmation, name-only connection screening, GPT fix reflection, and improvement recommendations). Three additional checks are confirmed from the research phase. Every check must show PASS before the report can proceed.
Here is a representative sample of what the gate checks:
Content and Sourcing
- Liability assessment completed and at LOW risk or below
- Every source URL fetched and confirmed — no hallucinated quotes or bios
- Review categorization (call/intake vs. program performance) included with statistics
- All external links archived on the Wayback Machine and swapped in the HTML
- No email addresses in the post — all redacted, contact form link only
- Zero first-person language anywhere in the report
Verification and Cross-Checking
- Corporate connections verified via state filings, not blog posts
- No name-only connections (similar names without shared principals or addresses)
- GPT review artifact exists with all Fix items resolved
- GPT findings are non-zero (proves the review actually happened)
- GPT fixes reflected in the final post
Completeness
- Entity tag box matches every named person and company in the report
- First-mention entity linking in HTML (every named entity linked on first reference)
- Facebook Ad Library checked and documented
- Reverse address search completed for all company addresses
- State registrations verified individually (actual database searches, not assumptions)
- “Correcting the Record” section present (debunking false third-party claims)
- Complaint response analysis with tone breakdown
- Review analysis completed with dual-classifier verification
- Inter-rater reliability computed (Cohen’s kappa)
- Results-Only Rating in report with methodology link
Liability Review
Before any report is uploaded to WordPress, it goes through a liability assessment. This is not about avoiding controversy — it is about ensuring every statement in the report is defensible.
The assessment checks for:
- Defamatory language: Does any statement accuse the company of illegal activity without sourcing to a court filing or enforcement action?
- Unsourced claims: Does every factual assertion have an inline source link?
- Opinion masquerading as fact: Words like “deceptive” or “misleading” are replaced with factual observations like “not disclosed” or “absent from website”
- Unfair framing: Are positive findings (clean enforcement record, strong BBB rating, improving trends) given equal prominence to concerning findings?
- Connection evidence: Is every corporate connection labeled with its evidence tier?
Reports must reach LOW risk before upload. If the assessment is MEDIUM or higher, the report goes back for revision until every identified issue is resolved.
The “Correcting the Record” Approach
Every report includes a section that corrects false claims made about the company by third-party websites. If a blog incorrectly states that Company A is owned by Company B, and my investigation found that the actual owner is Company C per state filings — that correction goes in the report.
Why This Matters: Correcting misinformation about a company — in the company’s favor — demonstrates that the report is fair and evidence-based, not adversarial. A company reading its own report should see that I investigated claims others made and set the record straight where warranted.
Report Philosophy
These reports exist to help consumers AND to encourage companies to do better. They are not adversarial. The philosophy:
- Praise what deserves praise. If a company has a clean enforcement record, strong complaint responses, or improving review trends — say so prominently.
- Frame gaps as opportunities. “Not disclosed on website” is something a company can fix tomorrow. I present it that way.
- Lead with strengths. Every section acknowledges what the company does well before noting what is absent.
- Let the data speak. Strong compliance + clean record + good responses = the data tells a positive story. Weak compliance + many complaints = the data tells that story too. I trust the reader to see it.
What This Process Does Not Do
Transparency means acknowledging limitations:
- It does not access private records. Sealed court filings, proprietary databases, private financial records, and internal company documents are not part of this analysis.
- It does not prove reviews are fake. A call/intake review is a real experience. The analysis separates review types — it does not invalidate them.
- It does not determine the “real” rating. The Program Performance Rating is one lens. Consumers who value customer service during intake should factor in those reviews too.
- AI classifiers have biases. They may over-index on certain keywords or miss context that a human reviewer would catch. The dual-classifier approach mitigates but does not eliminate this.
- Review samples are limited by platform access. Review platforms restrict how many reviews can be accessed without authentication. Samples of 30-50 reviews per platform provide directional insight but have wider margins of error than larger samples would. Important individual reviews may be missed.
- CFPB data is voluntary. Not all consumers file complaints, so CFPB numbers undercount actual issues.
- State registration searches have limits. Some state databases require CAPTCHAs or in-person access. When a state cannot be searched, the report says so.
- AI research may miss human context. Nuances from verbal communications, in-person interactions, or industry insider knowledge may not be captured.
- This is not legal, financial, or investment advice. It is a research compilation to help consumers start their own investigation.
Correction Policy
Standing Invitation: If any company believes its report contains errors — miscounted reviews, misclassified reviews, incorrect corporate filings, wrong complaint counts, or math errors — submit corrections with supporting data via the contact form. Documented corrections will be incorporated and noted in the report.
The classification criteria, sampling method, statistical calculations, and quality gate checks described on this page are fixed as of the publication date. They apply equally to every company reviewed. No company receives favorable or unfavorable treatment in how its report is researched or written.
Key Takeaways
- Up to eight specialized AI research agents independently investigate corporate records, court filings, consumer reviews, complaints, website compliance, marketing practices, nonprofit financial analysis (when applicable), and review classification
- Fifteen-plus public databases are searched — government records, corporate filings, consumer platforms, and web intelligence tools
- Review analysis collects 30-50 reviews per platform (limited by platform access restrictions), classifies each into three fixed categories, and uses independent dual-classifier verification measured with Cohen’s kappa
- The entire report is cross-checked by a second AI engine (GPT) for accuracy, consistency, and fairness
- A 43-point quality gate blocks publication until every check passes — no overrides, no exceptions
- A liability review ensures every claim is sourced, every connection is evidence-tiered, and the report is defensible
- False third-party claims about companies are corrected in the report — the process is fair, not adversarial
- Every limitation is disclosed — this process compiles public records, it does not claim to reveal hidden truths
Methodology version 1.3 — Updated February 2026. Changes from v1.2: Updated agent count from seven to eight (added conditional Agent 7 for nonprofit 990 analysis; renumbered Independent Review Classifier to Agent 8). Removed ConsumerAffairs from review platforms (not used in practice). Added PissedConsumer to data sources list. Added state registration verification process description (15+ state SOS databases searched in parallel). Updated quality gate from 24 to 43 checks reflecting automated script plus manual judgment phases. Added description of three-phase QA automation.
Frequently Asked Questions About How I Research Companies
How do I research the companies reviewed on this site?
Every report starts with up to eight specialized AI research agents working in parallel, each with a defined scope and specific databases to search. They don’t communicate with each other, specifically to avoid groupthink or one agent’s assumption contaminating another’s findings.
What sources does the research actually check?
The page names them specifically: state corporate filings (all 50 states where relevant), SEC EDGAR, OpenCorporates, FDIC BankFind, NMLS Consumer Access, CourtListener/RECAP and PACER for court records, the CFPB complaint database, BBB profiles, Trustpilot and Google reviews, and IRS Form 990 filings (via ProPublica) for nonprofits.
Why don’t I just tell you whether a company is good or bad?
Because facts change — complaint counts go up, leadership changes, and a company with a clean record last year may not have one today. Instead of a static opinion, the reports link you directly to the live public sources so you’re seeing current data, not a stale snapshot from whenever the report was written.
How is this different from other “company review” sites?
Two problems I built this to avoid: affiliate review sites that get paid to send leads, and AI-generated content that echoes unsourced assertions from other sites. My process requires every corporate-connection claim to be verified through actual state filings — claims that can’t be verified get flagged as unconfirmed rather than stated as fact.
How thorough is the process, in numbers?
Per the page: 8 AI research agents, 15+ public databases searched, 2 AI engines used to cross-check findings, and 43 quality-gate checks before a report is finished.