Development International e.V.

Methods

How do we know?

We choose research designs for the question at hand. This page sets out our toolkit by the kind of claim each design can support, its limits, and where we have applied it.

Levels of evidence

A research approach can be quantitative, qualitative or both. Within it, the research design sets the logic of inference, and so the kind of claim the findings can support. Data collection instruments serve the design. We group our designs by the strongest claim each supports in standard use.

Describe and estimateWhere do things stand, and how large is the problem?
Compare and trackIs it changing, and does the change hold?
ContributionHow and why did change happen, and what did the measure contribute?
AttributionWhat would have happened without the measure?

Contribution and attribution answer different questions. Theory-based evaluation, process tracing and QCA build the causal case through evidence on mechanisms or configurations of conditions; only a counterfactual comparison can attribute. Designs combine, and repeated waves or a comparison group move a study up the ladder. We state which claim a report makes.

Our research designs

Each entry states what it is, what to use it for, its limits, and where we have applied it. Designs marked CSDDD suit the evidence the EU Corporate Sustainability Due Diligence Directive asks of companies, including the periodic, indicator-based assessment of whether due diligence measures are effective.

Describe and estimate

Where do things stand, and how large is the problem?

Quantitative content analysis

analytic methodCSDDD

Systematic, replicable coding of a document corpus against a pre-specified framework. In our benchmarks the corpus is a population of corporate disclosures and the framework is derived from legal requirements and recognised standards, so scores can be compared across companies, sectors and years.

Applications
Regulatory benchmarking at scale; peer comparison; disclosure quality where consistency across companies matters.
Limits
Measures what companies disclose, not what they do. Scores are only as good as the codebook and the agreement between coders.
Pre-specified codebookKPI-based scoring frameworkDouble coding of a sample
  • LkSG HREDD Performance Assessment (2026): 141 practices, 25 effectiveness metrics
  • UK Modern Slavery Act: three studies (2018, 2019, 2022)
  • Dodd-Frank §1502 filing evaluations, reporting years 2014–2016
  • French Duty of Vigilance (134 companies, 42 KPIs); EU NFRD (Germany, Austria, Sweden); California TISC (n=1,504 and n=1,961)

Survey design

designCSDDD

Structured measurement of a defined population: cross-sectional surveys at one point in time, and population-based surveys with probability cluster sampling where a representative prevalence estimate is needed. Repeated waves or a panel move a survey up to the next rung.

Applications
Prevalence where no administrative baseline exists; what workers, communities and suppliers experience, know and trust.
Limits
Self-report and non-response bias; sampling error grows with design effects in clustered samples. Sensitive topics need indirect questioning (see instruments).
Structured questionnaireProbability cluster samplingWeighting for design effects
  • Workers in Fairtrade small producer organisations: 300 worker interviews across five countries
  • SBTi assessment: parallel surveys of staff (n=74) and Advisory Group members (n=82)
  • Cocoa sector oversight in Côte d’Ivoire and Ghana for the US Department of Labor (2008–2011): population-based surveys, 40+ clusters per country, evaluating compliance with the Harkin-Engel ProtocolDI associates

Econometric estimation and modelling

analytic method

Regression, extrapolation and equilibrium models that produce defensible estimates of phenomena that cannot be observed directly, with every assumption documented and confidence intervals reported.

Applications
Estimating the scale of hidden phenomena; projecting the effects of policy options.
Limits
Estimates inherit the assumptions of the model and the quality of the input data; they describe, they do not attribute.
Regression analysisExtrapolation from partial dataEquilibrium modelling
  • Estimating the number of hired workers in Fairtrade cocoa cooperatives, Côte d’Ivoire
  • Europe’s child labour footprint (€50 billion, 2019), using the Basu multiple equilibria model

Cost-benefit and cost-effectiveness analysis

decision-analytic methodCSDDD

What a measure costs per unit of outcome, where reallocating spend would do more good, and the economic costs and benefits of regulatory options. Built to be auditable and contestable.

Applications
Prioritising due diligence spend; value-for-money reviews; regulatory impact assessment.
Limits
Depends entirely on the outcome measures it is attached to and on how costs are attributed; results shift with the discount rate and the perspective chosen.
Costing frameworkCost-effectiveness ratiosValue-for-money criteria
  • Europe’s child labour import footprint: €50 billion estimate (2019)
  • Cost model of Dodd-Frank Section 1502 compliance (2011), submitted to the US Securities and Exchange CommissionDI associates

Value chain analysis

analytic frameworkCSDDD

A framework from the global value chains literature for mapping the actors, relationships, power and governance in a supply chain: its scope, who does what at what price and under what conditions, and how value and risk are distributed. Case studies, interviews and price data are gathered within it.

Applications
Mapping supply chains beyond tier 1; finding where governance gaps and leverage lie.
Limits
A framework for organising evidence, not a test of it; findings are descriptive unless combined with a comparative or causal design.
Interviews across chain tiersPrice and margin dataContract and audit review
  • Non-judicial grievance mechanisms: global mapping of the remedy ecosystem
  • Rainforest Alliance HREDD ecosystem study
  • Cocoa producer empowerment, western Côte d’Ivoire: 31 cooperatives and a 9-cooperative comparison group

Qualitative content and thematic analysis

analytic method

Systematic coding of media, policy and corporate texts into themes, developed inductively or from a deductive frame, following Braun and Clarke. Corpora are built by explicit rules, and AI-assisted extraction is validated by manual coding.

Applications
Public and stakeholder perceptions; how an organisation or controversy is framed; questions of reputation and legitimacy.
Limits
Interpretive by design: themes are patterns of meaning, not frequencies. Rigour rests on a transparent corpus, a documented coding frame and a second coder, not on statistical inference.
Rule-based corpus constructionThematic codingAI-assisted extraction, manually validatedSentiment analysis
  • SBTi assessment: 141 media articles from 58 outlets, with social media sentiment analysis

Compare and track

Is it changing, and does the change hold?

Time-series and cohort tracking

designCSDDD

Longitudinal observation: repeated benchmarks of the same population, cohorts of suppliers, sites or cases followed across cycles, and predictive validation that compares earlier risk scores with what happened afterwards.

Applications
Showing whether change happens and holds; testing whether a risk analysis predicts real incidents; tracking recurrence and severity.
Limits
Shows change, not cause: secular trends, changes in detection effort and attrition from the cohort can all masquerade as effects.
Repeated cross-sectionsLongitudinal cohortsPredictive validation
  • Annual Dodd-Frank §1502 benchmarks, reporting years 2014–2016, of the same population of filers
  • Repeated UK Modern Slavery Act benchmarks of the FTSE 100 (2018, 2019, 2022)

Linked buyer–supplier data analysis

designCSDDD

Procurement and transaction data joined with supplier outcomes over time, to test whether payment terms, price pressure or order volatility undermine what due diligence asks of suppliers.

Applications
Examining whether a company’s own purchasing practices contribute to the harms it seeks to prevent.
Limits
Observational: associations between buying terms and supplier outcomes can reflect supplier selection rather than buyer behaviour unless a comparison design is added.
Transaction data linkageSupplier panelsSupplier-side surveys
Applied in commissioned work not listed on this site.

Peer benchmarking

designCSDDD

One company’s results positioned against a reference population scored on the same indicators in the same period, using our published benchmark data.

Applications
Telling a board or supervisor where a company stands relative to its sector.
Limits
Cross-sectional: a position against peers says nothing about whether the company’s own measures caused it.
Published benchmark dataCommon indicator set
  • Reference population of 254 German companies across 16 sectors from the LkSG HREDD Performance Assessment

Contribution

How and why did change happen, and what did the measure contribute?

Case study

designCSDDD

An in-depth enquiry into a phenomenon in its real-world context, following Yin: single cases for depth, multiple cases for comparison and replication, and embedded designs with several units of analysis. Used to explain how and why something happened, including the root causes of recurring harms.

Applications
Mechanisms and root causes; settings too complex for measurement alone to explain an outcome.
Limits
Generalises analytically to theory, not statistically to a population. Rigour depends on case selection logic, triangulation and rival explanations being tested.
Key informant interviewsDocument reviewDirect observationTriangulation
  • Addressing root causes for UNICEF: three value chain case studies
  • Conflict minerals due diligence in telecoms: 12 companies
  • GIZ ABS Initiative evaluation: case studies in Côte d’Ivoire, South Africa and Kenya
  • Child labour monitoring and community dynamics, Ghana: embedded case study of two communities at two points in timeDI associates

Theory-based evaluation and contribution analysis

designCSDDD

Evaluation against an explicit theory of change, following Mayne: evidence gathered for each causal link, alternative explanations tested, and the strength of the contribution claim stated. Findings are structured against the OECD-DAC criteria where commissioners require it.

Applications
Building a credible causal case where no comparison group exists; showing a board or funder how measures lead to outcomes.
Limits
Establishes contribution, not attribution: it cannot say what would have happened without the measure, and it is only as strong as the theory of change and the evidence for each link.
Theory of changeContribution analysis protocolParticipatory validation
  • Incentive framework for artisanal gold mining, Dabakala, for Solidaridad (2024–25)
  • GIZ ABS Initiative evaluation (2025): ex-post design against OECD-DAC criteria
  • planetGOLD terminal evaluation for UNEP (under way)

Process tracing

designCSDDD

A within-case design that reconstructs the causal chain step by step, applying formal evidence tests to each link: hoop tests that a hypothesis must pass, and smoking-gun tests that confirm it.

Applications
Answering “did our measure cause it?” when there is only one case.
Limits
Inference is only as strong as the evidence tests and the rival hypotheses considered; a single case cannot show how typical the mechanism is.
Causal chain reconstructionFormal evidence tests
Applied in commissioned work not listed on this site.

Qualitative Comparative Analysis (QCA)

analytic methodCSDDD

A set-theoretic method for 10 to 50 cases, such as suppliers, sites or initiatives, that identifies which combinations of conditions are necessary or sufficient for an outcome, with crisp or fuzzy sets. Combinable with econometric analysis.

Applications
Portfolios where causes come in combinations.
Limits
Results are sensitive to how conditions are calibrated and to limited diversity among cases; configurations are explanatory patterns, not effect sizes.
Crisp-set and fuzzy-set QCACondition calibration
Applied in commissioned work not listed on this site.

Attribution

What would have happened without the measure?

Quasi-experimental designs

designCSDDD

Non-randomised comparisons that estimate what a measure caused: difference-in-differences, matched comparison groups, interrupted time series around the start of a measure, and stepped-wedge designs that use a staggered rollout as the comparison.

Applications
Attributing a change in outcomes to a measure, programme or regulation.
Limits
Each design rests on an assumption that must be tested and reported, and self-selection into a programme is the usual threat.
Key assumption
Parallel trends for difference-in-differences; selection on observables for matching; a stable pre-trend for interrupted time series; no time-varying confounding for stepped-wedge rollouts.
Difference-in-differencesMatchingInterrupted time seriesStepped-wedge rolloutFixed-effects panels
  • Impact evaluation of the OECD minerals Guidance (2023–25): difference-in-differences on mine-level data; enrolment in iTSCi was associated with an estimated 61% fewer violent incidents

Experimental designs

designCSDDD

Randomised or field-experimental comparisons where a measure can be assigned or a treatment tested directly: correspondence (audit) studies, vignette experiments embedded in surveys, and randomised rollouts where a programme is phased in anyway.

Applications
The strongest test of whether a practice, such as a recruitment or grievance procedure, treats people differently.
Limits
Rarely feasible for whole programmes; results answer the narrow question tested and may not transfer beyond the setting.
Key assumption
Random assignment holds and treatment does not spill over to the comparison group.
Correspondence studiesVignette experimentsRandomised phased rollout
Applied in commissioned work not listed on this site.

Published DI work   Work carried out by DI associates

What makes a finding trustworthy

The ladder above concerns internal validity: whether a causal claim is warranted. Four further tests apply to every study we produce.

Measurement validity

Indicators are derived from the legal or normative standard they are meant to reflect, and the codebook is published with the study. Disclosure scores measure what companies report. What happens in their supply chains needs field research, and we treat disclosure as its starting point.

Reliability

Benchmarks are coded against a pre-specified framework, a sample of documents is double-coded, and coder agreement is reported. Qualitative coding uses a documented frame and a second coder.

Representativeness

Benchmarks cover whole populations. Surveys use probability sampling where a prevalence estimate is the aim, with weighting for design effects and reporting of response rates. Case studies generalise to theory, and we say so.

Qualitative rigour and integration

Qualitative findings are judged by credibility, transferability and dependability: triangulation, documented case selection and an audit trail. In mixed-methods studies we state how the strands are combined, whether convergent, explanatory sequential or embedded.

Measuring due diligence effectiveness under the CSDDD

The EU Corporate Sustainability Due Diligence Directive, as amended in 2026, uses effectiveness in three senses. Measures must be adequate: capable of addressing adverse impacts and commensurate to their severity and likelihood (Art. 3(o), Art. 9). Their implementation and effect must be assessed at least every five years and after significant changes, using qualitative and quantitative indicators (Art. 15), and every prevention and corrective action plan must carry its own indicators for measuring improvement (Art. 10(2)(a), 11(3)(b)). And the Directive itself will be judged by its protective effect on rightsholders (Art. 36).

We derived 26 indicators from the obligations article by article. For each, the table gives what it measures, the question it asks, and a research design that could answer it.

IndicatorMeasuresWhat it asksDesignDesign type
C01Process qualityPredictive validity: share of actual adverse impacts that arose in areas the scoping exercise had flaggedRisk scores at the scoping date compared with impacts observed afterwards, per partner or areaprospective cohort · predictive validation
C02ProcessIn-depth assessment coverage: share of prioritised high-risk areas assessed within the cycleMapped coverage per cycle: share of spend, volume or partners in flagged areas covered by in-depth assessmentadministrative time series
C03AdequacyAlignment of effort with risk: resources and measures assigned to each impact against its severity and likelihood rankPortfolio analysis of measures and spend against the prioritisation registercross-sectional portfolio analysis
C04ProcessTime to the second tier: interval between addressing the most severe and likely impacts and starting on the restMilestone tracking across the prioritisation registeradministrative panel
11 of the 26 indicators reach outcomes for people and the environment, 4 measure an intermediate result, and 11 measure process, its quality or its adequacy. The Directive’s definition of appropriate measures pushes the balance toward outcomes.
Six indicators exist only because the law creates the obligation: alignment of effort with risk, transmission of assurances, SME support, the effect of industry initiatives, verification accuracy and responsible disengagement.
None of the designs is experimental in the strict sense, and almost every one needs data over time. Companies that start collecting baselines before July 2029 will be able to show change; those that start afterwards will only be able to describe a starting point.

Data collection instruments

Instruments serve a design.

Semi-structured interviewsPurposively selected informants; depth and context, triangulated with other sources
Structured questionnairesClosed-response items for comparable quantitative data
Indirect questioning for sensitive topicsList experiments and randomised response, so prevalence is estimated without exposing any respondent
Multiple systems estimationCombining incomplete lists to estimate the size of a hidden population
Document reviewCorporate reports, audits, programme files and policy papers
Participatory workshopsCo-design, validation and collective sensemaking with stakeholders
Direct observationStructured observation of sites, cooperatives and programme activities
Administrative and secondary dataIncident registers, audit records, transaction data, conflict event databases, trade statistics

Research ethics and data

Fieldwork rests on informed consent, anonymity and safeguarding, especially where children are involved. Studies involving people are submitted for ethics review where a commissioner or partner institution requires it, and we apply the same standards where none does. Personal data are handled under the GDPR, and data management plans set out what is collected, how it is stored and when it is destroyed.

Scoring frameworks are fixed before scoring begins and published with each study. Raw data are published where we hold the rights. We report what the evidence supports, including its limits. Our principles

Engraving of three researchers closely examining a sampling map
Research meeting, Abidjan

References

  • Beach, D. and Pedersen, R. B. (2019). Process-Tracing Methods: Foundations and Guidelines, 2nd ed. University of Michigan Press.
  • Braun, V. and Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101.
  • Creswell, J. W. and Creswell, J. D. (2018). Research Design: Qualitative, Quantitative, and Mixed Methods Approaches, 5th ed. SAGE.
  • Gereffi, G., Humphrey, J. and Sturgeon, T. (2005). The governance of global value chains. Review of International Political Economy, 12(1), 78–104.
  • Krippendorff, K. (2018). Content Analysis: An Introduction to Its Methodology, 4th ed. SAGE.
  • Lincoln, Y. S. and Guba, E. G. (1985). Naturalistic Inquiry. SAGE.
  • Mayne, J. (2012). Contribution analysis: Coming of age? Evaluation, 18(3), 270–280.
  • OECD DAC Network on Development Evaluation (2019). Better Criteria for Better Evaluation. OECD.
  • Ragin, C. C. (2008). Redesigning Social Inquiry: Fuzzy Sets and Beyond. University of Chicago Press.
  • Shadish, W. R., Cook, T. D. and Campbell, D. T. (2002). Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Houghton Mifflin.
  • Yin, R. K. (2018). Case Study Research and Applications: Design and Methods, 6th ed. SAGE.