Lending where there is no credit bureau
When most of your applicants have no formal credit history, the question isn't which scoring model to buy — it's which of your own data you've been throwing away.

Credit scoring literature assumes a bureau. It assumes that for any applicant you can retrieve a history of obligations and how they were met, and that your job is to model on top of it. In much of the market we serve, that assumption fails for the majority of applicants — not because they are bad credit risks, but because their entire financial life has happened outside the formal system.
The instinctive response is to look for a substitute dataset. Airtime top-ups, handset metadata, social graph, psychometrics. Some of this genuinely has signal. Most of it is oversold, much of it raises real fairness and consent problems, and almost none of it should be the first place you look.
The first place to look is what you already hold and are not using.
Most institutions have years of their own repayment history sitting in the core system, and are not learning from it in any structured way. They have deposit and transaction patterns for existing customers. They have data on which loan officers' recommendations perform and which don't. They have restructures, early settlements, and the outcome of every collections action ever taken. This is directly relevant, already consented to, already yours, and usually not being used beyond an arrears report.
Start there, and start simpler than you think you need to. A handful of well-chosen variables in a transparent model will beat an elaborate one you cannot explain, for three reasons that all matter more than a marginal gain in accuracy.
You can explain a decline. In most jurisdictions that is becoming a legal requirement, and in all of them it is how you keep a customer who was refused this year and is fundable next year. "The model said no" is not an answer anyone can act on.
You can detect when it stops working. Simple models degrade visibly. Complex ones degrade quietly, and in a volatile economy the relationships underneath will shift faster than your retraining cycle. A model fitted on a stable period will mislead through a currency shift, and you want that failure to be obvious rather than gradual.
And you can defend it. To a supervisor, to a board, to a customer. Explainability is not a nice-to-have in regulated lending; it is the thing that lets you use the model at all.
Two practical points that get missed constantly.
The first is that you can only learn from loans you approved. Your data contains no repayment history for anyone you declined, which means every model you fit is trained on a population your existing policy already selected. If your current rules are systematically wrong about a segment, the data will keep agreeing with them. The only way out is a deliberate, small, controlled volume of lending outside the current cut-off, treated as the cost of learning — and most institutions never do it, then wonder why their model recreates their assumptions.
The second is that in high-inflation and multi-currency environments, historical amounts are not comparable across time. An income figure from two years ago and one from today are different quantities wearing the same units. Anything amount-based has to be normalised — to the unit of account, to a basket, to a ratio of something stable — before it means anything in a model.
Finally: the loan officer's judgment is data. In institutions where officers know their borrowers, that assessment carries real predictive information, and the aim is not to replace it but to record it in a structured, comparable form so it can be tested. Some officers are much better than others. Currently, nobody knows which — because the assessment was never captured in a way that could be scored.
You probably do not have a data problem. You have a capture problem, and a not-looking-at-what-you-have problem.