Where Customers Come From
A visual read of the funnel, so the drop-offs are obvious without reading model metrics.
Restricted to TMA leadership + sales (nick@, shane@, shanemccormick@, dani@, marcos@thematchartist.com). Sign in with your TMA Google account.
Real-dollar view of what the open pipeline is worth today, where revenue is at risk, and which specific deals deserve attention this week. All numbers are model-derived from the v4 win-prob and cancellation-risk models, with median customer LTV applied.
Next-14-day meeting cohort, ranked by win-prob × median LTV. Cancel-prob shown in red where it's above 25%.
Day-of-week × hour close-rate heatmap on the booked-strategy-call cohort. Cells require n≥10. Darker green = higher close rate.
When the model says "30% chance to close", do 30% actually close? Diagonal = perfect calibration.
20 prospects who submitted info but never booked. The model ranks them by predicted conversion probability if re-contacted today.
Which acquisition channels produce real customers. $/lead is the expected revenue contribution of a lead from that channel (total LTV / leads acquired). Use this to decide where to spend the next $1.
Per-photographer net revenue, refund rate, and gallery yield. Minimum 10 shoots to appear. Recency = last 90 days.
Forecast of incoming payments based on customer plan classifications. Paythen $800/mo installments roll forward monthly until ~$3200 paid. Half-pay bookings (~$1400-$2000 down) have a balance due ~1 week after their HubSpot shoot date. Coaching revenue (Dani's date-coaching, tagged in Stripe) excluded.
Per-month mix of how new customers paid: full / half-pay / Paythen installment. Trend: Paythen + half-pay growth, full-pay decline through 2026.
Models tightened after a critical review pass. Most notable: Cox PH dropped sales_rep (post-customer field) → concordance fell honestly from 0.86 to 0.72. No-show classifier now trains on first-call-per-prospect only (removed per-email label duplication). Time-aware dedupes added. Win-prob no longer mislabeled as "leaky". Duplicate q_total_answer_words in lead-score features fixed. Booking-latency chart relabeled (it measures days-to-book, not rep response speed).
The 10 most actionable findings from all the models + data joins shipped to date. Ordered by impact. Each headline links to the full panel.
Things that go against conventional sales/marketing wisdom in our data.
Models that didn't work because of a data limitation. Surfaced so we know what to fix next.
Predictive models in production, ordered by how much money each one moves. AUC = ranking quality (0.5 = coin flip, 1.0 = perfect). Decile lift = top-10% concentration vs base rate. Click the chevron on any model for the decile-lift chart + top features.
What this is: P(close | this prospect just booked a strategy call). Trained on 2,462 historical booked prospects, 195 engineered features incl. pre-booking page views, email engagement, prior deal history, time-of-booking signals. What changed: v4 is the new generation — v2 (AUC 0.80) was retired; v4 hits AUC 0.860 with strictly pre-decision features and no leakage. Drives the Today tab Win % + Hot Calls expected-revenue ranking.
How to act on it: Top decile closes at ~96%, bottom decile at ~3%. Treat top-decile bookings as high-priority for rep prep; bottom-decile is where to invest screening time on the front end (or accept they'll churn).
What this is: P(prospect cancels or no-shows the booked call). 195 features, same engineering stack as win-prob v4. AUC 0.908 — the highest-quality model in the catalog. Trained on 2,462 historical bookings (38% baseline cancel rate). Drives the Today tab Cancel % column.
How to act on it: Top-decile prospects cancel ~97% of the time — proactive reach-out (text + email 24h before) is highest-leverage outreach work. Bottom-decile (under 3% cancel risk) is safe to leave alone.
What this is: P(eventually closes | just hit the site / submitted a form). Runs before booking — that's what distinguishes it from win-prob v4. Reports two cuts: all leads (incl. legacy backfill) and web-only (the cleaner real-traffic slice). How to act on it: Top-decile leads warrant same-day SMS + rep call; bottom-decile is a long-nurture cohort (email only).
What this is: P(this paying customer will refund before their shoot). Use this to intervene 1 week pre-shoot (extra check-in, stylist call, payment-plan adjustment). Top features: days between purchase & shoot, n_payments, sales rep, buyer_type, q_zero_dates_flag, app_none_flag. Decile spread: top 10% = 46% refund rate; bottom 10% = 1.4%. 33× lift. Bottom of page surfaces the at-risk active customers.
What this is: An earlier no-show model trained on lightweight call-time features (hour-of-day, day-of-week, host, prior calls). Now superseded by Cancellation Risk for prediction, but still surfaced because the bottom decile here flags "low-engagement" prospects (blank phone, unfilled questionnaire) — useful as a coarse triage filter even when cancellation risk is the operative signal.
What this is: Per-rep XGBoost classifiers. For each prospect, predict P(close | this rep). Recommend best rep at routing time. Per-rep training metrics + a historical analysis showing the close-rate lift if we'd actually routed everyone to the model's pick.
What this is: Per-photographer refund rates, smoothed using a Beta-Binomial prior so a photographer with 1 refund out of 5 shoots doesn't look like "20% refunds" when one bad-luck event would swing it back to 0%. How to act on it: Focus retention conversations on photographers whose entire 95% credible interval sits above the org-wide rate, not just whose point estimate happens to be high.
What this is: HDBSCAN unsupervised clustering on the five Claude-judge scores (intent / urgency / financial readiness / dating sophistication / effort) — groups customers into natural archetypes without telling the algorithm what to look for. How to act on it: Steer sales scripts toward the high-LTV / low-refund archetypes; opposite archetypes need better screening at booking time or careful expectation-setting on the call.
What this is: BG/NBD probabilistic model — given each customer's payment history, estimates P(transact again in next 365d). Caveat: Stripe records each captured payment as a separate transaction, so customers on multi-month payment plans inflate the repeat-rate. True repeat-shoots (4+ payments separated by months) are ~10% — those are the genuine re-engagement targets. How to act on it: Pair highest-P(alive) with longest gap-since-last-payment for the strongest re-engagement DM list.
What this is: Cox Proportional Hazards models time from lead-creation → customer, accounting for prospects who haven't bought yet (right-censoring). Concordance 0.72. How to act on it: Slow-decay segments are where extended nurture (SMS, retargeting) pays off; fast-decay segments warrant a same-week call.
What this is: For each rep, a 12-week rolling baseline of weekly call volume; any week that falls >2σ below baseline is flagged. How to act on it: A flag isn't proof of a problem — it's a prompt to check whether the rep was on PTO, switched roles, or actually slowed down.
What this is: BERTopic on the full Typeform free-text answers for two questionnaires we send: the Consultation Q (filled by 3,062 leads after they book a sales call) and the Pre-Shoot Q (filled by 1,928 customers after they purchase, before the shoot). Topics are forced via K-Means clustering on sentence-transformer embeddings; stopwords filtered.
How to read it: Each topic is named by its top n-grams. The Δ pp column shows prevalence in the positive outcome (became customer / refunded) minus the negative outcome.
What this is: BERTopic on Dani+Marcos transcripts cross-referenced against the real HubSpot lifecycle_stage = customer label. Joined via local CSV first, then live HubSpot API by E.164 phone + email for everything else. 100% match rate.
Top close signal: The "zero, card, perfect" topic (credit-card collection language) appears in 40% of customer calls vs 7% of non-customer calls. That's the closing-moment language pattern.
What this is: BERTopic discovered themes across 206 Dani/Marcos call transcripts (voicemails excluded). Each topic is then cross-referenced with the per-call AI-summary outcome (customer_paid / interested / declined / etc) to surface which conversation themes correlate with closes.
How to read it: The Δpp column = (% of positive-outcome calls that hit this topic) − (% of negative-outcome calls that hit it). Bigger positive Δ = topic concentrates in wins.
Sample-size caveat: Only 13 calls have a hard-negative outcome (declined + no_show); 61 are positive (customer_paid + interested + follow_up_scheduled). Loss signals are statistically thin — read this as "what wins talk about", not "what loses talk about".
What this is: A separate no-show model trained on the full HubSpot contact record — phone area code, derived state from desired shoot location, time-of-day, day-of-week, every questionnaire question + flag, plus a boolean for whether the questionnaire was even filled out.
How to read it: Feature importance below ranks signals that distinguish attendees from cancels. The decile lift table shows how concentrated no-shows are in the lowest-scored bucket — the bottom decile is roughly 100% no-show, the top decile roughly 0%.
Important caveat: The strongest signals are data-availability signals (blank phone, unfilled questionnaire). The model is mostly learning "do we have data on this person yet?" — operationally, that's still a useful triage: the lowest-decile prospects are the safest to deprioritize.
What this is: SMS is TMA's actual follow-up channel post-call. This bucket the prospect's outbound/inbound SMS volume and looks at close rate per bucket. Price-mention detection is a separate, binary feature.
How to act on it: Reps should escalate texting volume in the early-engagement bucket — that's where lift is largest.
What this is: There is no explicit "refund reason" field in HubSpot, so this panel infers reasons from contact_status + shoot_status_export values among refunded customers vs not.
How to read it: Pre-shoot cancellations dominate. Post-shoot dissatisfaction is rarer but more painful per occurrence.
How to act on it: Most refund mitigation lives in the post-booking / pre-shoot window. Stylist confirmations and shoot-brief PDFs target this exact window.
What this is: Customer count, revenue, and refund rate by geography. Toggle City ↔ State to switch rollup level. Toggle Table / Bubble / Heatmap for visualization. Use the timeframe selector to spot recency trends in specific markets.
How to act on it: Expanding markets = recent surge + lower-than-average refund rate. Refund hot spots warrant a closer look at the local photographer + travel logistics.
What this is: For each acquisition month, what percent of those customers had 2+ / 3+ / 4+ subsequent Stripe payments. Most "2+" is just multi-payment plans (i.e., not retention). True retention is at the 4+ column — those are people who came back for another full shoot package months later.
How to act on it: The 4+ column is the real signal for repeat business; older cohorts give the cleanest read since they've had time to mature.
What this is: For each pending lead, we compute the cosine similarity between their LLM-judge score vector (5 dimensions) and the average vector of top-LTV customers. Higher = "this prospect looks like our best customers on the dimensions we score."
How to read it: Five features is a thin similarity space — treat as a directional prioritization signal, not a verdict.
How to act on it: Cherry-pick these for next-day callbacks before SMS-volume work.
What this is: Among prospects who eventually book, this charts how long they took to book vs whether they ultimately closed. Important: This is NOT speed-to-lead. It's days from lead-creation to first booked call.
How to read it: >4 weeks shows 73% close (n=377), <1hr only has n=3 — too small to compare. Read this as "patient bookers close more often", not "respond slower."
How to act on it: Don't slow follow-up; do invest in long-tail nurture (SMS, retargeting) since slow-bookers are higher-value when they finally book.
What this is: Direct Google Search Console API pull, refreshed daily. 16-month rolling window for thematchartist.com. Replaces the previous manual CSV-export workflow with live data. Brand queries (anything matching "the match artist" / TMA) are split out so the non-brand view shows what SEO investment actually produces.
Filter: ≥10 impressions in both windows, position now worse than prior. Higher delta = bigger rank drop. Catch these BEFORE the traffic tanks.
Cities people are searching for that TMA either ranks poorly on or has no dedicated page yet. The geo expansion roadmap.
Each row is a TMA page's GSC clicks/position/CTR cross-referenced with how many of the people who first landed there became HubSpot customers.
What this is: Per-URL clicks from Google + revenue attributed to leads that landed on each URL. (Note: this section uses the older manual CSV export; the new Search Intelligence panel above uses live API data.)
The headline: City pages convert 335× better per click than blog pages, despite blog driving 94% of organic traffic.
How to act on it: SEO investment should heavily skew toward city-page coverage, not blog volume. The top blog pages still bring leads, but $/click is dwarfed by city-page ROI.
Web traffic (Plausible), HubSpot funnel events, and Stripe payments — monthly, 2023-02 → present. The big story: traffic peaked in 2024 and the visitor → lead conversion has compressed.
This is the business story in one row: how many people become leads, how many book, how many show, and how many actually become customers.
A visual read of the funnel, so the drop-offs are obvious without reading model metrics.
How the booked-lead universe currently breaks into follow-up types.
The strongest positive and negative trigger families after stripping out leakage and noisy sales-process fields.
The buyer types that close best after they show up.
Question themes from the questionnaire. This is useful for understanding intent, not just scoring.
High-level traits that keep showing up among stronger buyers.
The most important friction patterns coming out of the queue logic, text analysis, and false-positive audits.
When leads tend to book better. This is safe, top-of-funnel timing, not downstream sales-process noise.
These are the clearest operational buckets right now. Each one comes with real names behind it in the explorer below.
The acquisition and behavior combinations that keep appearing around stronger leads.
Monthly cohorts help keep the story honest when recent leads have not had time to mature yet.
| Cohort | Leads | Booked | Q / Booked | Showed / Booked | Close / Showed |
|---|
Predicting net spend per paying customer at first-payment time. Useful for ranking — TMA's tier distribution is narrow, so this is package-tier prediction (basic vs premium), not whale-finding.
Customers ranked by predicted LTV, bucketed into 10 deciles. Each bar is the average actual net spend in that decile. The wider the spread, the more useful the model is for prioritization.
From all 2,121 paying customers. The top decile predicts $2,726 mean actual spend vs $1,856 in the bottom decile (1.47×).
| Name | Predicted | Actual | Decile | Buyer |
|---|
Joins the new Stripe customer data to the HubSpot funnel by email. Where does our $4.79M actually come from — and which segments quietly leak refunds?
Reset buyers (post-divorce / long-relationship) lead at $2,527 ARPC. Ready buyers, surprisingly, have the lowest ARPC AND highest refund rate.
The pre-call question they asked. Process / profile-help themes monetize highest; price-question theme still pays full tier ($2,260 ARPC).
Refund-rate spread across card brands is small but real. Missing-brand rows are likely ACH or alt-payment paths.
Where the money actually comes from. State and city ranking, refund hotspots, and the long tail of countries.
Net revenue per state (top 24). Bar length = revenue, dot = ARPC. Click a state to filter the Lead Explorer.
States ranked by refund rate (min n=15). Where customers regret the most.
Mostly US, but the international tail.
From Stripe card metadata at customer creation.
When booked-call activity happens, year-over-year traffic, cumulative revenue.
Each cell = number of strategy-session bookings in that hour-of-week slot. Reds darker where booking activity peaks.
Plausible monthly visitors overlaid by year. The 2024 → 2025 collapse is the headline.
Net revenue accumulating over the dataset window.
Refunded $ as a % of gross monthly revenue. Spikes flag bad-shoot months.
Booking rate around US holidays (top 12 windows by lift).
Stage-by-stage drop-off, monthly cohort close rates, time-from-lead-to-customer.
From the visitors who hit the site to the customers who paid. Bars scaled to the largest stage.
Monthly cohorts × close rate. Rows are creation months, color intensity is the % that became customers.
Days from lead creation → booked sales call, for customers only. Faster = stronger intent.
Days-since-creation for non-customers still in the funnel. Long tail = stale queue.
Honest AUC of the show / show_q / close / close_q models.
How spend, payment patterns, and customer ranks distribute.
Net spend per paying customer, 24 bins from $0 to $4,500. The bimodal shape is real (basic vs premium tier).
Each dot is a held-out validation customer. Diagonal = perfect prediction.
Live ranking from Stripe + LTV scores.
Of paying customers, how many made one payment vs spread across installments?
Where the $281K of refunds is concentrated. Use to prioritize delivery quality fixes.
From the refund-risk model. Low AUC, but the tail is the directional flag.
Which segment costs the most in refunds (absolute dollars).
Refund rate per Stripe card brand. Missing-brand bucket likely ACH or alt-pay.
Credit / debit / prepaid split among paying customers.
Buyer-type bubble, app preferences, goal types, self-description distribution.
Bubble size = n customers, x = ARPC, y = refund rate. Top-left = ideal segment (high ARPC, low refund).
From the pre-call questionnaire.
Hinge / Bumble / Tinder / multi-app split.
Marriage / serious / casual / unspecified.
Which pre-call question themes monetize best after the call.
From the Plausible export — what pages drive entrances, what events convert, where traffic comes from.
Most-entranced URLs across the full window. The Tinder/Hinge/city pages dominate top of funnel.
Tracked events. Submitted-Contact-Form is the dominant lead capture, "Booked A Call" the bottom-of-funnel.
Aggregated visitor source. Google + Direct dominate.
Predicting who will refund using only what's knowable at first payment.
LightGBM gain. The model still ranks customers — just don't trust the absolute probability when AUC is near random.
Customers who haven't refunded yet, ranked by model score. Useful as a triage queue, not as a calibrated probability.
| Name | Risk | Spend | Stage | Buyer |
|---|
Use this to move from the big picture into actual people. The default modes are intentionally plain-English.
The core model metrics and top feature summaries are still here, just moved out of the main operating view.
Full generated text reports behind the dashboard.
Per-photographer lifetime revenue, refund rate, and average net per customer. Filter by individual photographer + time range — all-time, presets, custom dates, or canonical year/quarter/month buckets.
Cities ranked by lifetime revenue and refund rate. Austin and San Francisco dominate volume; smaller markets (Minneapolis, Miami) show highest mean LTV.
City pages (/{city}-dating-photography and /online-dating-photographer-near-me) are TMA's only meaningful organic lead source —
blog posts (Tinder / Hinge / Bumble) get 10-20× the traffic but convert at ~0.05% vs city pages at ~10-16%.
Three stacked views of the city / blog / home / service / other buckets. Time window applies to all three charts.
Side-by-side: city-page Google clicks (GSC, top-of-funnel) vs every HubSpot lead created (right axis) vs every strategy call booked (right axis). If clicks fall but leads/calls hold, attribution is broken or other channels are filling in. If all three fall together, top-of-funnel is the bottleneck.
Raw Google clicks landing on each bucket per month. Top-of-funnel reality.
HubSpot contacts with hs_analytics_first_url matching each bucket (solid lines) plus the TOTAL leads across all sources (dashed white line, includes the ~25% of contacts with no first_url). The gap between green (city-attributed) and white (total) shows the unattributed pool.
Strategy calls booked (pclose opp_dt) attributed to each bucket via the contact's first_url. Dashed white line = total calls (all sources, ~99.8% attributable).
Only computable for GSC-covered months. Blog stays near zero — it's a brand-awareness funnel, not a lead funnel.
The long view. Plausible visitor totals go back to Feb 2023 and all-HubSpot lead-creation goes back to Sep 2023 — both predate GSC coverage. Useful for spotting the absolute peak of TMA traffic and where the funnel actually drifted from.
Side-by-side snapshot of each bucket: how clicks and leads moved between the two most recent quarters in the GSC-covered window.
Pick any subset of city pages to overlay their click + lead + conversion histories. Click a chip to toggle. Default = top 5 by total clicks. Search to add cities not in the default list.
All city pages with GSC data. Click a column header to sort. Click ▸ on a row to expand the full monthly chart for that city (clicks · leads · calls · L/100C). Pick a window below — the table aggregates dynamically.
Your original hypothesis: "City pages aren't getting an insanely less amount of views as they used to, but our leads are down significantly… they just aren't converting to calls like they used to." This panel finally tests it directly — Plausible page-views (anyone landing on the page, any source), GSC clicks (Google-search arrivals only), and HubSpot leads attributed to that page, monthly, for every city page that has GSC traffic.
Did views drop more, the same, or less than leads? "Views holding while leads collapsed" would mean a conversion problem. "Views and leads dropping together" means the bottleneck is upstream (deliverability / SEO / referral).
The original prospect-funnel hypothesis: someone reading "why aren't I getting matches on Tinder" gets told they need better photos, googles
"cincinnati online dating photographer" or "dating app photographer denver", and lands on the corresponding money page.
This table confirms exactly that pattern — the dominant search intent on every city page is some variant of dating app photographer {city} or
dating profile photographer {city}. Average position 8-15 (page 1-2 of Google), CTR 1-4%.
Click ▸ to expand a row and see the top 15 queries for that city.
City pages where the lead-per-100-click rate dropped the most in the last 3 months vs the prior 3 months. Filter: ≥30 clicks in both windows. Empty / near-empty list = conversion is broadly holding or improving (the current state of TMA).
Top 30 individual landing pages (hs_analytics_first_url) ranked by HubSpot lead count over the full HubSpot history (2023-09 onwards),
with close-rate to customer. This is "what's actually working", independent of GSC ranking.
Six top-line findings from 233 blasts + 9,068 subscribers + $4.45M lifetime revenue. Click any card to jump to that section.
233 sent broadcasts via Drip from 2021-01 → 2026-03, paired into 119 originals + 107 resend-to-unopened follow-ups. Every blast joined to daily HubSpot leads, strategy-call bookings, and Stripe revenue. The questions: are blasts moving the business, which ones, and what's the deliverability decay costing us?
Open rate by year, weighted by recipients. The drop from ~29% (2023) to ~8% (early 2026) is the most likely culprit behind the lead-volume softness in 2025-26. Google's bulk-sender rules tightened in Feb 2024; combined with list age, the path-to-spam-folder is paved.
OLS distributed-lag regression on 1,842 days of daily revenue vs lagged opens (1-7d, 7-14d, 14-28d, 28-90d windows). Coefficients show the dollar value attributable to each opens-window. Sum across all lags ≈ $7,269 in 90-day revenue per 1,000 opens. R² is low (0.038) — opens are one of many revenue drivers — but the AVERAGE effect direction is consistent.
Monthly aggregate of opens (blue), strategy calls (red, since 2024-04), HubSpot leads (gold dashed, since 2023-09), and Stripe net revenue (green, full history). Visually inspect whether peaks in opens precede peaks in calls/revenue.
9,074 active Drip subscribers joined to $4.9M Stripe lifetime revenue (1,854 are paying customers). Scored on
R (Recency, days since signup — newer = higher),
F (Frequency, Drip's auto-computed engagement lead_score),
M (Monetary, lifetime $ from Stripe joined by email).
Tags are how Drip auto-segments subscribers. Total monetary by tag tells you which segments are worth emailing — and which (looking at you, "Dead") are dragging deliverability down.
Drip captures landing_url + original_referrer at subscribe time. Mean monetary per acquisition source reveals which pages and referrers bring high-LTV subscribers. (Filtered to n≥10 subscribers per row.)
Days from subscriber created_at → first Stripe payment, filtered to subscribers who paid AFTER signing up (avoids Drip's auto-tag-customers-post-purchase confound).
If most converters do it fast, Drip is a post-conversion engagement tool, not pre-conversion warming.
Each bubble is one blast. X-axis = send date. Y-axis = 30-day revenue lift ($). Bubble size = recipients. Bubble color = open rate (green high → red low). Hover for subject. Vertical clusters reveal cadence patterns; the color trend from left to right shows deliverability decay; upper-envelope bubbles are the candidates worth subject-line-cloning.
GitHub-style daily grid for the last 2 years. Border-dot on a cell = a blast was sent that day. Cell color intensity = daily Stripe net revenue (relative to the 2-year max). Visually reveals send-rhythm voids, seasonal patterns, and post-send revenue "warmth" that visually correlates with sends.
The 25 blasts with the biggest revenue jump in the 30 days after send vs the 30 days before. Some of these reflect real causal lift; others are noise from random timing. Pattern to watch: story-driven subjects ("James found the girl...", "47 Year Old Dad Goes on Unlimited Dates...") dominate the top.
Each row pairs an original blast with its "(resend to unopened)" follow-up. Combined reach = parent recipients + resend recipients (disjoint by construction since resend only goes to non-openers of the parent). Resends typically yield 2-7% open rate vs originals at 9-38%, but they capture real incremental opens — about 30-40% of total combined opens come from the resend.
Full list of 233 sent blasts. Click column headers to sort. Click ▸ for full-content details (subject, send time, recipient count, opens %, clicks %, attributed revenue lift, etc.).
Each cell is a (day-of-week, hour) bucket showing send count, mean open-rate %, and mean 30-day revenue lift across blasts sent at that slot. The 7×24 grid reveals which timing slots historically produced the highest engagement and downstream lift.
For every blast, we capture the average daily count of calls / leads / revenue at offsets −30 to +90 days from the send. The lines show the average over all blasts (200+ per offset). The vertical line at day 0 is send-day. If blasts cause a real lift, you'd see the lines bump up at and after day 0.
Blasts sent within ±X days of a US holiday (or "none" for the rest) — comparing mean open % and mean 30-day revenue lift across holiday categories. Watch for outliers (Christmas, Valentine's, Memorial Day).
Pearson correlation between per-blast metrics and 30-day revenue lift (since 2024-11 when click data became reliable, n≈68). If clicks correlate stronger than opens, prioritize click-rate as the deliverability canary — that's also what the deliverability research recommends (Apple MPP pollutes opens).
Three angles on "when does a blast matter most":
Splits blasts into lead-state quartiles based on HubSpot leads in the week before send. Tests whether blasts boost lift more when the funnel is hot or cold.
Most TMA blasts are paired with a resend within 7 days, so "solo" is rare. Difference here measures whether resends add or cannibalize.
Bigger sends to larger pools — do they produce proportionally more lift, or does signal saturate?
Days bucketed by quartile of rolling-7-day opens preceding that day — i.e., days where TMA had high recent email engagement vs days with low. Mean daily revenue, calls, leads per quartile. If higher recent-opens days have higher outcomes, that's dose-response evidence.
Univariate Pearson correlations between subject-line features and (a) open-rate %, (b) 30-day revenue lift (originals only — resends muddy interpretation). Heavy ML model (LightGBM + sentence-transformer embeddings + SHAP) is training on Mac mini in background — will refresh this section once results land.
Monthly correlation between blast volume (sends, opens) and monthly outcomes (revenue, calls, leads). Note: these correlations are confounded by time-trend — e.g., as deliverability decayed in 2024-26 both opens AND leads declined, inflating their correlation. Still useful as a sanity check.
Every email a subscriber receives, mapped by their funnel path. Combines all 28 Drip campaigns (full email sequences, freshly pulled) with the 15 HubSpot workflow campaigns observed in the event log. Split by booked-a-call vs never-booked cohorts so you can see "this is what people get sent in each path".
The Win Prob v5 model relies heavily on questionnaire-text PCA components (50%+ of total gain). Here's the decomposition: TF-IDF logistic regression on the same text identifies the actual words that predict convert vs non-convert, plus a 3× conversion-rate spread across quintiles of the dominant text direction.
v5 of the win-prob model rebuilt with Phase 4 email-engagement + questionnaire-text features. Same target as v4 (P(close) for booked calls), trained on 2,462 historical booked calls. v2 of cancellation risk rebuilt on customers-only with the same Phase 4 features.
LightGBM classifier trained on "is this subscriber a top-LTV customer" (top quartile, ≥$2,799). Combined with cosine-distance to the 5 nearest top-LTV customers in scalar-feature space. Top 1,000 non-customer lookalikes exported as CSV — upload to Meta/Google as custom audience seed.
Aggregated win/loss analysis on 107 Dani+Marcos call subjects (28 became Stripe customers, 79 didn't). Two layers: AI-extracted objections + buying signals from the per-call summaries, and raw bi/tri-gram n-gram counts across transcripts. Small-sample caveat applies — strongest signals are the most over/under-represented phrases.
MiniLM sentence embeddings of every subscriber's HubSpot questionnaire + Typeform consultation free-text answers, UMAP-reduced and HDBSCAN-clustered. Reveals 6 distinct subscriber archetypes with wildly different conversion rates.
Cox PH regression + Kaplan-Meier curves on the journey from Drip signup → first paid charge. 17.6% of converters do it within 30 days. After that, conversion curve is nearly flat — so the first 30 days of subscriber attention is when the action happens.
Per-state TMA performance + per-MSA breakdown joined to the TMA Marriage Map dataset (346 MSAs · ACS PUMS · BEA RPP · religion · politics · migration). The under-penetration table identifies metros with huge target-pool (single men 28-44, high median income) where TMA has barely any presence.
Apples-to-apples comparison of three LightGBM classifiers with progressively richer features. Same target (is_customer), same time-ordered 5-fold CV, same n=9,087 subscribers.
Quantifies exactly how much predictive lift email engagement adds, and gives a dollar-attribution of email signal to TMA's revenue.
Reverse-causation guard + outcome-leakage guards applied throughout (see Method).
Joined every per-customer data source on email: Drip CSV, HubSpot contacts (incl. questionnaire free-text + photographer + shoot status), Stripe payments & customers, HubSpot email events (automated Workflow + Meetings emails, the thing you asked about), Typeform consultations, and JustCall logs. For every converted customer, we computed the full email journey to first paid charge.
Live-pulled from Drip: how many subscribers actually opened the recent blasts, what Google measures, and the research-grounded plan to climb back out of Promotions/Spam. The thesis you proposed is correct: concentrating sends on the engaged cohort produces measurable open-rate recovery on the full list over 4-8 weeks.
The next-level analytics from joining 9,068 Drip subscribers to HubSpot first-touch metadata + Stripe lifetime revenue. 6 pre-signup customers excluded (reverse-causation guard — Drip auto-tags subscribers "Customer" the moment a Stripe charge clears, which back-fills profiles). Everything below uses the post-filter dataset of 9,068 subscribers → 1,848 paying customers (20.4%) → $4.45M lifetime revenue.
LightGBM classifier predicting P(convert), multiplied by median customer lifetime ($2,295) to estimate pLTV.
Top 200 actionable non-customer subscribers (excludes Dead-tagged, capped to last 3 years of signups).
This is the deliverable for Dani — sorted by predicted lifetime value.
Three model variants trained; AUC + decile lift below the table give the honest read on signal strength.
Each row is one Drip signup month, showing how that cohort progressed.
Conversion % = fraction of the cohort that eventually paid.
$/sub = total cohort revenue ÷ cohort size (the true unit economics of acquiring a subscriber that month).
Days-to-customer = median time from signup → first paid charge (only counted for converters who paid AFTER signing up).
Cohorts with n < 5 excluded as noisy.
Most recent month is incomplete — short maturity period explains low conversion.
Six funnel stages from Drip subscription → Stripe purchase. Each bar's width is proportional to its share of subscribers. The biggest single drop is the bottleneck. Click any stage band for definition.
Where the highest-LTV subscribers come from. Rows = landing-page family, columns = referrer domain, cell color intensity = mean $/subscriber.
Cells show n subscribers, conv%, and mean $/sub. (Drip landing_url falls back to HubSpot hs_analytics_first_url when missing.)
Reads the same way as the Conversion tab's per-city table — but here the columns are referrers, revealing channel-page interactions.
Anderl-Becker removal-effect Markov chain over subscriber journeys. For each funnel stage and each acquisition channel, we compute the drop in conversion probability if that node is removed from the graph. The resulting attribution share is causally cleaner than first-touch/last-touch because it accounts for which states are actually necessary for conversion.
Per blast, we compute the lift in non-email signals (Google clicks, HubSpot leads by source, Stripe revenue, Typeform consultations) in the 7-day window after send vs the 7-day window before. If blasts produce brand-halo lift (people search Google for TMA after seeing the email, or arrive via direct traffic), we'd expect a positive median lift across 233 sends. Wilcoxon signed-rank test, one-sided alternative='greater'; reject H₀ at p < 0.05.
This is a forward-going experiment design — not yet running. The cleanest path to a causal lift estimate is to randomly hold out 10% of recipients from each blast and compare per-arm 30-day outcomes. Below is the protocol, sample-size math for 80% power, and key caveats. Numbers are computed from TMA's actual baseline conversion rate and per-blast revenue variance.
From the research agents — what to do about the open-rate collapse:
Next-level analyses worth implementing once we know the basics:
Cross-references Amex Uber charges (~/Downloads/activity.csv) with scheduled shoot dates from HubSpot. Each Uber charge is assigned to the NEAREST shoot within ±2 days (deduplicated — a charge can't double-count across overlapping shoot windows). Mike Schmidt is owner id 702762280.
Every Schmidt shoot that had Uber charges within ±2 days, with the charges themselves. Click ▸ to expand each shoot's individual rides.
ZIP codes extracted from Uber's Extended Details column. Schmidt's home ≈ 10010 (Gramercy/Flatiron Manhattan). NYC airports = LGA (11371), JFK (11430/11434), EWR (07114). Cross-referenced with Schmidt shoot dates to identify travel-day patterns (departure day-of, return day-after).
Charges that don't correlate to any photographer's scheduled shoot. Could be personal trips, photographer pre/post-shoot travel beyond 2 days, or other expenses.
Combined analysis of activity -2.csv (2025 full year, 1,772 txns) + activity.csv (2026 YTD, 327 txns).
Categorization, subscription detection (only consistent-amount, regular-cadence merchants), shoot-charge correlation, and Uber-by-photographer attribution using city + date matching.
Custom categorization based on merchant keywords (Amex's own categories are alongside for cross-reference).
Merchants with 2+ charges of consistent amount (CV < 15%) on regular cadence (weekly/monthly/quarterly/annual). Airlines / hotels / rideshare filtered out. Annualized estimate is what each subscription costs you per year at current cadence.
Each Uber charge attributed to the nearest photographer's shoot, prioritizing matches where the Uber's city = shoot city (high-confidence) over date-only matches (lower-confidence).
Non-Uber charges within ±3 days of any photographer's shoot — flights, hotels, meals, gas, parking, etc. Useful for verifying which trips were business vs personal.
Every TMA Amex airline ticket parsed from Extended Details (passenger, route,
departure day, ticket #), clustered into physical trips by photographer using
home-airport boundary detection (a trip starts when leaving home and ends when returning).
Each trip is matched to the shoots on that photographer's Google Calendar that fall inside
the trip window, and to hotel/car/rideshare charges paid on the same photographer's Amex
card during the same window.
Every shoot in the Travel tab is cross-referenced against the HubSpot client export, Stripe payment history, the Airtable Schedule-Raw export, and the Amex receipt geographic data. Score 4 = all four sources agree; lower = something doesn't line up.
Total Amex airfare + ground spend, matched-trip count, and combined cost per shoot for each photographer in scope.
How far in advance the outbound ticket was booked relative to departure day, across every clustered trip. The 5–6 week target window noted in the Apr 2026 pricing pivot is the right column to grow — anything ≤ 1 week is the dominant today.
Each dot = one airline segment. Looks for the "buy later, pay more" pattern. Hover for details.
Air vs. ground over time. Travel doubled in early 2026 per the pricing pivot — verify on actuals here.
Where the travel dollars flow. Counts include side-hops within a multi-city trip; shoots column shows how many shoots that destination produced.
Trips where total Amex spend (airfare + ground) per matched shoot was highest. Either had very few shoots, very high cost, or both. Use as a watch-list.
Efficient tours — multiple shoots packed into a single trip, keeping per-shoot cost low.
Every clustered trip. Click ▸ to expand each one for segment-level breakdown, matched shoots, and ground charges.
Live snapshot of what's happening right now — calls booked, closes attributed, and revenue captured this week. Refreshes every 5 minutes.
Each rep's week-to-date: calls on the books, calls completed, attributed closes, and contracted revenue. "Booked $" = sum of contract prices from this week's closes (not yet collected — includes Paythen / half-pay forward exposure).
Upcoming sales calls in the next 48 hours per rep, with our lead-score and no-show-risk attached. Sort by no-show risk to flag who needs a personal reminder. Click a rep heading to filter.
Three side-by-side rankings: total calls, attributed closes, and close rate. Honors the Rep + Range filters below. Bars scale to the leader in each column.
All reps (or filtered rep). Strategy calls vs total. Closed-to-customer overlaid where attribution exists.
Unique closes (bars, left axis) and close rate (line, right axis) per period, for all reps or the selected rep above. Close rate = unique closed prospects ÷ sales calls, matching the rest of this tab.
Monthly sales calls vs all new paying customers that month (Stripe truth). Since Dani started, the only sales reps are Dani & Marcos, so every close in this window is credited to the team — unlike the per-rep chart above, which counts only closes tied to a logged strategy session. Gold % = closes ÷ sales calls.
Zoomed view of the two newer reps, scaled to their volumes. Lets you see Dani vs Marcos head-to-head without Nate's volume flattening the chart.
Every strategy session, follow-up, and pre-shoot call pulled from each rep's Google Calendar back to 2022. Close attribution joins prospect emails captured on the calendar invite to the customer outcome in live HubSpot.
Sales calls = strategy + follow-up only. Pre-shoot column is shown for context but it's operational (photographer prep with existing customers), not sales — Marcos's 107 pre-shoot calls are from his photographer role, not him selling. Close rate denominator = sales-attempt calls where the prospect email was on the invite (~15% of older calls; ~100% of newer Calendly-routed calls).
One line per rep showing monthly strategy-call volume. Click a legend item to isolate a rep.
Heatmap of call activity by day-of-week × hour (America/Chicago local time). Darker = more calls. Honors the Rep filter. Useful for capacity planning + spotting underused hours.
For the selected rep (or all reps combined), the path from total calls → strategy/follow-up sales attempts → unique prospects → attributed closes.
Running total of closes per rep over time. Steeper slope = faster closing. Dani & Marcos only (current sellers; Nate, Nick & Shane excluded — their old Calendly events didn't capture prospect emails so attribution is unreliable).
Per-rep customer count, revenue, and refund rate using the HubSpot Sales Rep field at the customer level. Showing the current sales team (Dani and Marcos).
Donut of strategy / follow-up / pre-shoot / other for the selected rep. Tells you whether a rep is mostly running discovery (strategy) or mostly nurture (follow-up).
Read before drawing conclusions.
Photographer roster, upcoming shoots, on-the-books sales calls, and travel availability — pulled from the photographer-booking Firestore (sales.thematchartist.co) on the last pipeline run. Updates nightly.
Currently bookable. Lifetime stats joined from the Stripe + Airtable + HubSpot data.
Per-photographer busy/free calendar pulled from their Google Calendars every 30min (via sales.thematchartist.co Firestore). Red = busy, green = open. Use this to book shoots without double-booking. Click a column header to filter.
Last 30 days + next 120 days. Filter by photographer; toggle calendar / map / table.
Current week schedule per rep, pulled live from sales.thematchartist.co Firestore.
Calendar of recent + upcoming sales calls. Host = sales rep running the call. Kind: strategy = initial discovery, follow-up = post-call decision.
Live web-analytics for every thematchartist.co subdomain and the main Webflow site. Pageviews, CTA clicks, slider drags, video plays, form submits, lead identifications, session replays, and ad-hoc HogQL — pulling from project 404240.
Where thematchartist.com ranks on Google for the city "dating photographer" terms we deliberately target — pulled daily from keyword.com (counterpart to the GSC "Search Intelligence" panel: GSC = what Google shows us; this = where we rank). Updated nightly.
Per-city coverage — keywords tracked, how many rank in the top 3 / page 1, and average position.