Linear attribution gives every touchpoint equal credit whether it mattered or not. Time-decay attribution assumes recent touchpoints matter more, whether or not the data actually supports that. Every rule-based attribution model shares this flaw: the credit-assignment formula is decided in advance, before anyone looks at what the data actually shows. Algorithmic, data-driven multi-touch attribution models – built on Markov chains or Shapley values – flip that order, deriving credit directly from observed conversion path patterns rather than imposing a fixed rule on top of them.
This guide covers how the two dominant algorithmic attribution methods actually work mathematically, when each one makes sense, what neither can honestly tell you, and a practical path for building custom attribution infrastructure for paid campaigns rather than relying on a platform’s built-in model alone.
Rule-Based vs. Algorithmic Attribution: The Fundamental Difference
Rule-based models – first-touch, last-touch, linear, time-decay, position-based – apply the same fixed credit-assignment logic to every conversion path, regardless of what actually happened in that specific journey. This makes them simple and completely transparent, but structurally arbitrary: a linear model splits credit evenly across five touchpoints whether all five genuinely contributed or one did all the real work.
Algorithmic, or data-driven, models instead analyse actual observed conversion paths – which sequences of touchpoints led to a conversion, which didn’t – and derive credit assignment from patterns in that data itself. The two methods that dominate this space, Markov chains and Shapley values, take genuinely different mathematical approaches to the same underlying question, and understanding both is necessary before deciding which one, or which combination, fits a specific paid media programme.
The Markov Chain Model: Attribution via Removal Effect
A Markov chain models the customer journey as a graph, where each touchpoint (an ad channel, an email open, a website visit) is a node, and the transition probability between nodes reflects how often customers actually move from one touchpoint to the next in the observed data, ultimately arriving at either a conversion or a non-conversion state. Because it’s built on transition probabilities, a Markov chain captures the sequential order of interactions, not just which touchpoints appeared somewhere in a path.
Credit for each channel is calculated using the removal effect: the model computes the probability of conversion across all observed paths as normal, then recalculates that probability with a specific channel entirely removed from the graph. The difference between those two numbers – how much conversion probability drops when a channel is taken out of the picture – becomes that channel’s attributed contribution. A channel whose removal barely changes the overall conversion rate gets little credit; a channel whose removal causes a steep drop gets substantially more.
This approach has two practical advantages for paid media teams. It scales more easily than the alternative described below, since the underlying calculation doesn’t grow combinatorially as more channels are added. And it’s easier to explain to a non-technical stakeholder, since “what happens to conversions if we take this channel away” is an intuitive framing that doesn’t require explaining game theory.
The Shapley Value Model: Attribution via Game Theory
Shapley value attribution originates from cooperative game theory and calculates each channel’s marginal contribution by considering every possible ordering in which the channels in a conversion path could have occurred, not just the specific order that actually happened. For each possible ordering, the model measures how much a given channel adds when it’s included versus when it’s left out – its marginal contribution in that specific ordering – and the final Shapley value for a channel is the average of that marginal contribution across every ordering considered.
This exhaustive approach is what makes Shapley value the more theoretically “fair” of the two methods: rather than relying only on the paths that actually occurred, it evaluates every combinatorial possibility a channel could appear in, producing a more complete picture of genuine incremental contribution. The cost of that completeness is computational – the number of orderings to evaluate grows factorially with the number of channels involved, meaning Shapley value calculations become expensive quickly once a model is tracking more than a handful of distinct channels, and can become impractical without simplifying the channel taxonomy first.
Markov vs. Shapley: Choosing Between Them
| Factor | Markov Chain | Shapley Value |
| Theoretical basis | Probability theory; transition states and removal effect | Cooperative game theory; marginal contribution across all orderings |
| Computational cost | Scales reasonably with more channels | Grows factorially; expensive beyond a handful of channels |
| Interpretability for stakeholders | Intuitive via “what happens if we remove this channel” | More abstract; harder to explain without a game theory analogy |
| Sensitivity to click spam | Lower | Higher – numerous low-value interactions can distort results more |
| Best suited for | Larger channel sets, teams needing to explain results to non-technical stakeholders | Smaller, well-defined channel sets where maximum theoretical fairness matters most |
Neither model is universally superior – the choice depends on how many distinct channels are being modelled and how much the audience for the results values mathematical completeness versus intuitive explainability.
What Neither Model Can Honestly Tell You: Correlation vs. Causation
This is the caveat most attribution content undersells, and it applies to both algorithmic methods equally: Markov chain and Shapley value models both derive credit from observed conversion paths, which means they’re fundamentally analysing correlation, not causation. A touchpoint appearing frequently in converting paths may genuinely have influenced the outcome, or it may simply be a common step people take regardless of whether it changes their eventual decision.
This limitation has a specific, practical consequence for paid media: neither algorithmic method reliably detects conversion hijacking – cases like brand-term bidding or affiliate coupon codes capturing credit for a conversion that would have happened anyway, regardless of that specific touchpoint’s presence. A conversion lift or incrementality test, which deliberately withholds exposure from a control group, is the only way to distinguish genuine causal contribution from a channel that simply shows up reliably in the data without actually driving the outcome. Comparing algorithmic attribution results against periodic incrementality testing – rather than trusting either in isolation – gives a far more honest picture than either method provides alone.
Building a Custom Attribution Model: A Practical Path
- Start with GA4’s built-in data-driven attribution before building anything custom. For most organisations, this remains the single highest-return first step, since it requires no custom infrastructure and immediately reveals the gap between current channel spend and what a data-driven model would recommend. Custom Markov chain or Shapley value infrastructure represents the next level of sophistication, worth pursuing once GA4’s built-in model has been thoroughly used and its limitations genuinely become a constraint.
- Assess data readiness before committing to custom infrastructure. Both algorithmic methods need a genuine volume of conversion path data – touchpoint sequences leading to both conversions and non-conversions – exported from GA4’s BigQuery export or collected through server-side tracking. A dataset too thin to show stable patterns will produce a custom model no more reliable than a simpler rule-based approach, at considerably more implementation cost.
- Choose an implementation approach matched to available technical resources. Open-source tools exist specifically for this – the Marketing-Attribution-Models Python package (developed by DP6) implements both Shapley value and Markov chain models directly against a conversion-path dataset, offering a reasonable starting point rather than building the underlying mathematics entirely from scratch.
- Handle scale deliberately. Large conversion-path datasets can cause serious memory and processing issues when computing a full Markov chain model directly. A practical technique – grouping identical journey paths and working from path-occurrence counts rather than one row per individual customer journey – preserves all the information the model needs while dramatically reducing the data volume actually processed.
- Validate the custom model against both rule-based attribution and GA4’s data-driven model before trusting its output for real budget decisions. Large, unexplained divergence between methods is worth investigating before acting on it, rather than assuming the newest or most sophisticated model is automatically the most correct one.
- Layer in incrementality testing on a recurring basis, treating it as the causal check against the correlational picture either algorithmic model provides on its own.
- Build a genuine reallocation framework, not just a reporting dashboard. A custom attribution model that reveals a spending imbalance but never actually changes budget allocation delivers no more business value than the rule-based model it replaced – treat reallocation as an ongoing, tested process, not a one-time reaction to a single model run.
A Simplified Markov Chain Removal Effect Calculation
The following illustrates the core removal-effect logic in Python, using a small, illustrative conversion-path dataset rather than a production-scale implementation:
import pandas as pd
from collections import defaultdict
# Example: each row is a customer journey and whether it converted
paths = pd.DataFrame({
“journey”: [
“Paid Search > Email > Conversion”,
“Social > Paid Search > Conversion”,
“Paid Search > Conversion”,
“Social > Email > No Conversion”,
“Email > Conversion”,
]
})
def parse_path(path_str):
steps = path_str.split(” > “)
converted = steps[-1] == “Conversion”
return steps[:-1], converted
def total_conversion_rate(journeys):
conversions = sum(1 for _, converted in journeys if converted)
return conversions / len(journeys)
journeys = [parse_path(p) for p in paths[“journey”]]
baseline_rate = total_conversion_rate(journeys)
channels = set(ch for steps, _ in journeys for ch in steps)
removal_effects = {}
for channel in channels:
filtered = [(steps, conv) for steps, conv in journeys if channel not in steps]
if filtered:
rate_without = total_conversion_rate(filtered)
else:
rate_without = 0
removal_effects[channel] = baseline_rate – rate_without
print(removal_effects)
This simplified version illustrates the underlying logic – recalculating conversion rate with each channel excluded – that a production Markov chain implementation formalises using proper transition-probability matrices across a genuinely large dataset, ideally through an established package rather than a hand-rolled version at scale.
Common Mistakes When Building Custom Attribution
- Jumping straight to Markov or Shapley value modelling before establishing a GA4 data-driven attribution baseline. Custom infrastructure is a meaningful investment that’s only worth making once a simpler built-in model has genuinely been exhausted.
- Ignoring memory and processing constraints at scale. Running a full Markov chain calculation against an ungrouped, row-per-journey dataset can become computationally impractical well before a dataset reaches genuinely large scale; frequency-based journey grouping solves this directly.
- Treating algorithmic attribution output as proof of causation. Both methods analyse correlation in observed paths; neither can distinguish a channel that genuinely drove a conversion from one that simply appears reliably alongside conversions that would have happened regardless.
- Ignoring conversion hijacking risk. Brand-term bidding and coupon or affiliate codes can accumulate credit in an algorithmic model without genuinely driving incremental conversions – a gap only incrementality testing reliably closes.
- Applying Shapley value to too many channels without simplifying the taxonomy first. The factorial growth in computational cost makes this impractical past a certain channel count; grouping related channels into broader categories before modelling keeps the calculation tractable.
- Building the model but never changing budget allocation based on it. A sophisticated custom attribution model that never informs an actual spending decision provides no more practical value than the simpler model it replaced.
Frequently Asked Questions
What’s the difference between Markov chain and Shapley value attribution? Markov chain attribution calculates each channel’s contribution using the “removal effect” – how much conversion probability drops when that channel is excluded from the customer journey graph. Shapley value attribution, from cooperative game theory, calculates a channel’s average marginal contribution across every possible ordering of channels in a journey, which is more theoretically complete but computationally more expensive.
Do I need a data science team to build a custom attribution model? Not necessarily from scratch. Open-source Python packages implementing both Markov chain and Shapley value models already exist, meaning a team with reasonable Python and data-handling skills can implement custom algorithmic attribution without building the underlying mathematics entirely from first principles.
Should I use algorithmic attribution instead of GA4’s built-in data-driven attribution? Not necessarily as a replacement. GA4’s built-in data-driven attribution is typically the highest-return starting point for most organisations, and custom Markov chain or Shapley value infrastructure is best considered a further step once that built-in model’s limitations become a genuine constraint on decision-making.
Can algorithmic attribution detect fraudulent or low-value conversions, like brand bidding? No, not reliably. Both Markov chain and Shapley value models work from observed conversion path correlations, not controlled experiments, so they can’t distinguish a channel that genuinely drove a conversion from one that simply captured credit for a conversion that would have happened anyway. Incrementality or conversion lift testing is the appropriate tool for that specific question.
How much conversion data do I need before building a custom attribution model? There’s no fixed universal threshold, but both methods need enough conversion path volume to reveal stable, meaningful patterns rather than noise. A dataset too thin to show consistent patterns will produce a custom model no more reliable than a simpler rule-based approach.
Is Shapley value always more accurate than Markov chain attribution? Not necessarily in practice, even though it’s more theoretically complete. Its computational cost grows factorially with the number of channels, meaning it can become impractical for larger channel sets, where a well-implemented Markov chain model may deliver a more usable and still meaningfully more accurate result than a rule-based alternative.
The Bottom Line
Custom algorithmic attribution – built on Markov chains, Shapley values, or both – genuinely improves on the arbitrary, fixed logic of rule-based models by deriving credit directly from observed conversion patterns. But both methods remain fundamentally correlational, neither detects conversion hijacking on its own, and building custom infrastructure is only worth the investment once a simpler baseline like GA4’s data-driven attribution has been genuinely exhausted. Pair whichever algorithmic method fits your channel complexity with periodic incrementality testing, and make sure the resulting insights actually change budget allocation rather than sitting in a dashboard. Search Savvy’s performance marketing services and Google Ads and PPC services build attribution and budget-reallocation strategy around exactly this kind of data-driven approach, and the performance marketing glossary is a useful reference for the terminology covered throughout this guide.





