6 Ways Clubs Use Player Similarity Models in Scouting and Recruitment
A player similarity model is a statistical tool that compares footballers by their performance profiles and ranks how closely one player's numbers resemble another's. Recruitment departments use it to widen searches, test assumptions and find alternatives to expensive targets. The raw inputs are the match and player statistics recorded by data providers and platforms such as RubiScore (https://rubiscore.com), turned into comparable profiles.
This listicle covers how such models are built in broad terms and six practical ways clubs use them, followed by the limits every user should keep in mind.
How a Similarity Model Works
Most similarity models follow the same basic steps, even when the details differ between clubs and analysts.
First, a set of metrics is chosen to describe a player's style and output: for example, passes into the final third, progressive carries, shots, touches in the box, tackles and interceptions. Second, those metrics are converted to rates, usually per 90 minutes, so that players with different playing time can be compared. Third, the rates are standardised, often into percentiles or scores relative to other players in the same position, so that no single metric dominates because of its scale.
Finally, each player becomes a point in a multi-dimensional space, and the model measures the distance between points. Common techniques include Euclidean distance and cosine similarity, which compares the shape of two profiles rather than their overall size. The closest points are the most similar players.
The choice of metrics is the most important decision. A model built on attacking outputs will find similar attackers; one built on passing and positioning data will find players with a similar role in possession. There is no neutral version.
1. Finding Replacements for Departing Players
The most common use is succession planning. When a key player is likely to leave, whether through a transfer, contract expiry or age, the club needs to know who could fill the same role.
A similarity model takes the departing player's profile and returns a list of players whose numbers look most alike. That list is rarely the final answer, but it gives scouts a data-driven starting point that covers far more leagues and players than any scouting network could watch in person. It can also reveal that a player's contribution is unusual, with few close matches anywhere, which is itself valuable information for planning.
2. Widening the Search Beyond Familiar Markets
Scouting departments naturally concentrate on the leagues they know best. Similarity models help break that habit by searching across every competition in the dataset at once.
A player from a smaller league may produce a profile very close to a target in a major league at a fraction of the cost. The model will not say whether that player can handle the step up, but it flags candidates who would otherwise never be considered. This is where league-strength adjustments become essential, because raw numbers in weaker competitions tend to look stronger than they would at a higher level.
3. Defining Roles More Precisely
Positions on a team sheet are broad labels. Two central midfielders can play completely different roles, one a deep playmaker and the other a box-to-box runner. Similarity models help clubs describe roles through data rather than labels.
By clustering players with similar profiles, analysts can identify groups that behave alike regardless of listed position. That clarifies what a coach is actually asking for. A request for "a new number eight" becomes a request for a player whose profile sits in a specific cluster, which narrows the search and improves communication between recruitment staff and coaching staff.
4. Testing Scouting Opinions
Scouts form strong impressions from watching matches, and those impressions are valuable. Similarity models provide a second opinion.
If a scout believes a young player resembles an established star, the model can test whether the numbers support the comparison. Sometimes they do; sometimes they show that the resemblance is visual rather than statistical, based on a style of movement or appearance rather than output. Neither source is automatically right, but disagreement between them is a signal to look more closely.
5. Benchmarking Value and Contracts
Similarity models can support negotiation and valuation. If a target's profile is close to several players who moved for known fees or earn known salaries, those comparisons provide a reference point.
This use requires caution. Transfer fees and wages depend on age, contract length, the selling club's position and market conditions, not just performance. A similar profile does not mean a similar price. Still, clubs use comparable-player analysis as one input among many, much as property valuations use comparable sales.
6. Tracking Development Over Time
Similarity models are not only about other players. They can compare a player with his own past profile, showing how his style is evolving.
A young winger whose profile is drifting towards that of an inside forward, or a full-back whose numbers increasingly resemble a midfielder's, may be developing in a direction that suits a different role. Academies and loan departments use this kind of analysis to guide positional decisions and to judge whether a loan spell is producing the intended development. Game-by-game records, such as the appearance and position histories available on RubiScore, add useful context about where and how often a player actually featured.
Which Metrics Usually Go Into the Model?
The metric set depends on the role being searched for, but most models draw from a few families:
- Ball progression. Progressive passes, progressive carries and passes into the final third, which describe how a player moves the ball forward.
- Chance creation. Key passes, expected assists and passes into the penalty area.
- Finishing and shooting. Shots, non-penalty expected goals and touches in the box.
- Defensive activity. Tackles, interceptions, pressures and ball recoveries, often adjusted for the opponent's possession.
- Involvement and security. Touches, pass completion and turnovers, which describe how often a player is used and how safely he keeps the ball.
- Physical and aerial output. Aerial duels and, where tracking data exists, distance covered and sprint counts.
Analysts often weight these families differently for each role. A search for a defensive midfielder might emphasise ball recoveries and pass security, while a search for a winger might emphasise carries, dribbles and chance creation. Getting the weighting right is often more important than the similarity formula itself.
Common Mistakes When Reading the Output
A list of "most similar players" looks authoritative, which makes it easy to misuse. Frequent errors include treating the top name as the recommended signing, ignoring the similarity score gap between first and tenth place, comparing players across positions without checking whether the model was built for that, and forgetting that the output only reflects the seasons included in the data.
The Limits Every User Should Know
Similarity models are powerful but narrow. Several limits apply to every version:
- Context is flattened. A player's numbers reflect his team's style, his coach's instructions and the quality of his teammates. Two identical profiles can come from very different situations.
- Metric choice shapes results. Change the inputs and the list of similar players changes, sometimes dramatically.
- Sample size matters. Profiles built on a few hundred minutes are unreliable. Most analysts set a minimum threshold before including a player.
- League strength is hard to adjust. Adjustments help, but no method perfectly translates performance between competitions.
- Intangibles are missing. Personality, adaptability, injury history, language and willingness to relocate do not appear in the data.
- Similar does not mean equal. A player can match a profile in shape while operating at a lower level of quality.
How the Model Fits Into a Recruitment Process
The best recruitment processes treat similarity models as a filter, not a decision. A typical workflow might run as follows:
- Define the role and the metrics that describe it.
- Run the model to generate a broad list of statistically similar players.
- Apply hard constraints such as age, contract status, budget and eligibility.
- Review video of the shortlisted players to check whether the numbers match what happens on the pitch.
- Send scouts to watch the most promising candidates live.
- Combine data, video and scouting reports in a final assessment.
At every step, human judgement adds context the model cannot capture.
A Tool for Better Questions
Player similarity models do not find the perfect signing automatically. What they do well is widen the pool of candidates, sharpen the definition of a role, challenge assumptions and provide reference points for value. Used alongside video analysis and traditional scouting, they help clubs ask better questions about players, and make fewer decisions based on reputation alone.