Credit scoring
Credit scoring is a statistical method that turns borrower data into a number predicting the likelihood of repayment, used to decide, price and limit loans.
Credit scoring is the use of a statistical model to convert information about a borrower into a single number that estimates how likely they are to repay. The score is compared against a cut-off, and the cut-off drives the decision to approve, refer or decline.
The claim underneath scoring is narrow and worth stating precisely: characteristics that predicted repayment among past borrowers will predict it among future ones. A score is not a judgement about a person. It is a ranking device β it says this applicant resembles a group that repaid at a certain rate, nothing more. It carries no opinion about any individual and makes no promise about any individual outcome.
That narrowness is also the source of its power. A model applies the same rule to every applicant, in milliseconds, at effectively zero marginal cost, and it does not get tired, targeted or persuaded. Where the volume of loans is high and the value of each is low, no form of human loan appraisal competes on cost.
Score, scorecard, rating and report
Four terms that are used interchangeably and should not be.
- Credit score: The number produced for a particular applicant at a particular moment
- Scorecard: The model itself β the characteristics, their weights and the points structure that generate scores
- Credit report: The factual record of a borrower's accounts and repayment history, held by a bureau
- Credit rating: An assessment of an entity's creditworthiness, usually expressed on a letter scale and applied to companies and governments rather than individuals
- Probability of default (PD): The estimated likelihood of default that a score maps to, expressed as a percentage
A score is derived from a report by a scorecard. The report is evidence; the score is inference. A borrower disputing a decision is usually disputing something in the report, and the two need to be separable to answer them.
How a scorecard works
Most scorecards in production are points-based models derived from logistic regression. Each characteristic is divided into bands, each band carries points, and the points are summed.
The construction sequence, in outline:
- Define good and bad. A "bad" is usually a loan reaching a specified arrears status β commonly 90 days past due β within a defined window. This definition determines everything downstream, and changing it later invalidates the model.
- Set the observation and performance windows. Characteristics are taken as at application; performance is measured over the following period, typically 12 to 18 months. The model can only be built once enough loans have had time to go bad.
- Assemble a development sample. Enough loans, and critically enough bad loans, to estimate reliably.
- Select and band characteristics. Test each candidate variable for predictive power and stability, then group its values into bands with materially different bad rates.
- Fit the model and allocate points. Convert the fitted coefficients into a whole-number points scale that staff and systems can work with.
- Validate on a holdout sample the model has not seen.
- Set the cut-off and the actions attached to each score band.
- Monitor and recalibrate as the applicant population and the environment change.
Worked example: a simple application scorecard
Time in business
- Under 1 year: 0 points
- 1β3 years: 15 points
- Over 3 years: 25 points
Borrowing history with lender
- First cycle: 0 points
- 2ndβ3rd cycle: 20 points
- 4th cycle or more: 30 points
Credit bureau result
- Adverse record: Auto-decline
- Clean, thin file: 10 points
- Clean, established file: 20 points
Debt-to-income ratio
- Above 45%: 0 points
- 35β45%: 10 points
- Below 35%: 20 points
Savings or deposit behaviour
- None: 0 points
- Irregular: 5 points
- Regular: 15 points
An applicant two years in business, on their third loan cycle, with a clean but thin bureau file, a DTI of 38% and regular savings scores 15 + 20 + 10 + 10 + 15 = 70.
- 75 and above: Automatic approval at full requested amount
- 60β74: Approve, with amount capped or additional security required
- 45β59: Refer for manual appraisal
- Below 45: Decline
At 70, the applicant is approved with a capped amount. The referral band matters more than it looks: it is where a small lender keeps human judgement in the process for the cases where the model is least confident, rather than discarding judgement entirely.
Types of scoring
Application scoring decides on new applicants using information available at application. This is what most people mean by credit scoring.
Behavioural scoring uses the borrower's own conduct on existing accounts β repayment timing, utilisation, balance trends β to decide on limit increases, top-ups, renewals and pricing. It is generally more predictive than application scoring, because past behaviour with you beats inferred behaviour from a population. Any lender with repeat borrowers can build behavioural scoring long before it has the data for a full application scorecard.
Collection scoring ranks delinquent accounts by likelihood of recovery, so that effort goes where it produces the most money. A lender with more arrears cases than officers is implicitly prioritising already; scoring makes the prioritisation explicit.
Bureau scoring is produced by the credit bureau itself across all its reporting members, so it sees obligations to other lenders that an internal model cannot.
Custom or internal scorecards are built on the lender's own portfolio. They fit the actual customer base far better than a generic model, and they require the lender to have a portfolio worth modelling.
Alternative data scoring
Conventional scoring assumes a bureau file. In markets where most borrowers have never held a formal loan, that assumption excludes the majority of the population, and alternative data is the route around it.
- Mobile money transaction history: Income regularity, balance patterns, expenditure discipline, business turnover
- Airtime and data purchase behaviour: Income stability and cash-flow smoothness at very high frequency
- Handset and telco tenure: Stability, contactability, length of verifiable footprint
- Utility and rent payment records: Willingness to meet recurring obligations
- Merchant or POS transaction data: Verified business turnover and seasonality
- Psychometric assessment: Attitudes and traits associated with repayment where no financial history exists
- Group and savings-scheme records: Contribution discipline and peer standing
Mobile money history is the strongest of these in practice, because it is transactional rather than declarative, hard to fabricate, and dense enough to reveal seasonality. It is also the most sensitive: it is a complete record of a person's financial life, and using it requires a lawful basis, informed consent and defensible retention rules under the data protection law of each market you operate in. The compliance question here is not a formality β it is frequently the binding constraint on whether an alternative data model can be deployed at all.
Scoring and judgement
- Cost per decision: Near zero (Scoring) vs. High (Judgemental appraisal)
- Speed: Seconds (Scoring) vs. Hours to days (Judgemental appraisal)
- Consistency: Complete (Scoring) vs. Varies by officer and by day (Judgemental appraisal)
- Thin-file borrowers: Handled poorly (Scoring) vs. Handled well (Judgemental appraisal)
- Unusual or complex cases: Handled poorly (Scoring) vs. Handled well (Judgemental appraisal)
- Auditability of the rule applied: Total (Scoring) vs. Partial (Judgemental appraisal)
- Data required to build: Substantial (Scoring) vs. None (Judgemental appraisal)
The realistic structure for most lenders is not one or the other. Scoring handles the volume β small loans, repeat borrowers, standard products β and appraisal handles the referrals, the first-time larger borrowers and the exceptions. Deciding which loans go down which path is a policy choice, and it is usually a better first project than trying to score everything.
What to do before you have enough data
A lender with 800 loans and 40 defaults cannot build a statistically valid scorecard, and building one anyway produces a model that encodes noise with the appearance of rigour. The sequence that works:
- Write the policy rules down first. Hard knock-outs and eligibility criteria, applied consistently. This is not scoring, but it delivers most of the consistency benefit immediately.
- Capture the data you will need later. Every characteristic you might want to model has to be recorded at application, in structured fields, from now β including on the applications you decline. A model built only on approved loans is blind to the applicants you rejected, and that gap is not recoverable retrospectively.
- Define bad and record arrears status consistently so that performance is measurable when the time comes.
- Start with behavioural scoring on repeat borrowers, where the data is your own and arrives fastest.
- Build the application scorecard once there is a sufficient body of matured loans with genuine bad outcomes.
- Recalibrate, and keep recalibrating.
Cut-off strategy
The cut-off, not the model, determines the business outcome. Moving it trades volume against loss, and the trade is the actual decision.
Illustrative figures for one portfolio:
- Cut-off 50: 78% approval rate, 9.5% bad rate on approved
- Cut-off 60: 64% approval rate, 6.2% bad rate on approved
- Cut-off 70: 47% approval rate, 3.8% bad rate on approved
- Cut-off 80: 29% approval rate, 2.1% bad rate on approved
There is no correct row. The right cut-off is the one that maximises profit given the product's margin, the cost of an acquired customer, the cost of collections and the lender's appetite β and a high-margin, short-tenor product can rationally accept a bad rate that would ruin a thin-margin secured one.
The swap set is the useful lens when changing a cut-off or replacing a model: not how many applicants are affected, but which ones. A new model that approves the same number of loans while swapping out a slice of bad accounts for a slice of good ones has improved the portfolio without changing volume at all.
Measuring whether a scorecard works
Discrimination β the model's ability to separate good from bad β is usually reported as the Gini coefficient, KS statistic or AUC. Higher is better, all three move together, and what counts as acceptable depends heavily on the portfolio and the data available; a thin-file microloan scorecard will not reach the discrimination of a mature secured-lending model, and should not be judged against one.
Two failures are more common than poor discrimination at build time:
Population drift. The applicants arriving today differ from the development sample β new channel, new product, new geography, new marketing. Population stability monitoring catches this, and it is the most frequent reason a scorecard that validated well degrades in production.
Calibration drift. The ranking still works but the loss rates attached to each band have moved, usually because conditions have changed. The scorecard is still sorting correctly while the cut-off has silently become wrong.
Both are found by monitoring, and monitoring is the part of scoring that gets built last and abandoned first.
Risks and limitations
Proxy discrimination. A model that excludes protected characteristics can still discriminate through variables correlated with them β location, handset type, name patterns, employer. This happens without anyone intending it, and it is only detected by testing outcomes across groups deliberately.
Explainability and adverse action. Many jurisdictions require that a declined applicant be told why. A points-based scorecard can answer this directly by identifying which characteristics cost the most points. More complex machine learning models cannot always do so without additional work, which is a live regulatory constraint rather than a technical inconvenience.
Self-fulfilling data. Applicants you decline generate no performance data, so the model never learns whether it was right about them, and the population it sees narrows over time.
Gaming. Once the rule is known, it is optimised against β by borrowers, by agents, and by staff with volume targets. Characteristics that are cheap to manipulate degrade fastest.
Structural breaks. A model built on pre-shock data does not describe a post-shock population. Currency movements, sector collapse, policy changes and pandemics all break the assumption that the past predicts the future.
False precision. A score of 68 looks more exact than "reasonably creditworthy". It is not. The underlying estimate has a confidence interval nobody displays.
Where credit scoring goes wrong in practice
Buying a generic model and applying it unchanged. A scorecard developed on a different population, product and market is a starting hypothesis, not a decision rule.
Building on approved loans only. Reject inference exists precisely because this bias is severe and invisible.
Never recalibrating. A scorecard is a photograph of a population at a moment. Left alone for four years it becomes decoration.
Overriding constantly. If officers override a large share of decisions, either the model is wrong or the overrides are, and nobody is tracking which. Override rate and override performance are the two numbers that answer it.
Scoring without a data protection basis. Alternative data models in particular are built on personal information that requires lawful grounds to collect, process and retain. Retrofitting compliance after deployment is far harder than designing for it.
Confusing the score with the assessment. The score estimates willingness and likelihood. It does not verify identity, confirm the loan's purpose, establish the value or enforceability of collateral, or check that the security documents are real.