Illustrative SCC Nexus visual representing vaping and public-health evidence

Health evidence · research infrastructure · public observatory

SCC Vaping

An independent, evidence-led research and public-information project designed to continuously collect, cross-reference, test and communicate evidence relating to vaping.

SCC Vaping is the active project identity. The former internal name Vaping26 remains in some filenames, workflow names and data contracts for compatibility and historical provenance, but it no longer defines the project or imposes an artificial 2026 end date.

Official short synopsis

Evidence before conclusion.

SCC Vaping is an independent, continuously updated evidence observatory that combines deterministic data engineering, systematic evidence surveillance, human review and carefully bounded machine learning to collect, cross-reference and audit research relating to vaping. Its private research engine preserves provenance and applies a deny-by-default publication firewall; a separate public observatory communicates only validated, publication-safe material. The project is also testing whether modern ML and source-grounded AI can reduce evidence-review workload while maintaining high recall, transparency, reproducibility and human scientific accountability. Its purpose is not to advocate a predetermined position on vaping, but to make the evidence base easier to interrogate, verify, challenge and update.

Principal aim

Not to prove that vaping is good or bad.

The central aim is to build an auditable system capable of asking increasingly precise questions and showing what the available evidence can — and cannot — support.

  • Continuously identify relevant studies, trials and official evidence.
  • Detect missing, duplicated, corrected or retracted evidence.
  • Keep smoking-cessation evidence separate from never-smoker, youth and exposure questions.
  • Retain provenance down to source records where possible.
  • Test competing explanations and expose uncertainty or disagreement.
  • Publish only material that passes validation and publication controls.
  • Test whether modern automation, ML and AI can make evidence surveillance substantially faster without making it less trustworthy.

Three deliberately separated layers

Research machinery, public communication and academic collaboration.

01 · Private research engine

sccvaping/research

The scientific machinery. Its pipeline is register → collect → snapshot → validate → cross-reference → normalise → audit → aggregate → publication firewall.

Current sources include PubMed, Europe PMC, ClinicalTrials.gov, Crossref/Retraction Watch, MHRA, Companies House discovery, ONS/NHS prevalence data, OHID Fingertips and optional OpenAlex enrichment. Raw and investigative material remains private; public release is deny-by-default.

ML may prioritise review, detect anomalies or assist bounded tasks, but may not independently establish causation, illegality, evidence grades or synthesis eligibility. AI extraction must remain source-grounded, provenance-preserving and human-reviewed.

02 · Public observatory

SCC Vaping Observatory

The public repository contains a deployed branded evidence website covering evidence, health, cessation, young people, prevalence, methodology, review and academic collaboration, together with public-safe research-status data, evidence summaries and a regulatory timeline.

The next step is to evolve this from a conventional website into an interactive evidence observatory with searchable studies, filters, timelines, trial/publication relationships, provenance drill-down and downloadable public-safe datasets.

Open the public observatory ↗
03 · Academic workspace

Google Drive collaboration layer

A separate academic workspace provides curated sharing areas for Studies and Evidence, Reports, PDF and Paper Library, Methods and Protocols, Data and Exports, Presentations and Correspondence, and Archive.

This provides an appropriate collaboration surface for academics who need more than the public website without exposing the private working engine.

Evidence status

Provenance · integrity · verification · validity

Full evidence record →
P

Provenance

Evidence intake distinguishes study records, protocols and post-publication notices, retaining source and study-design context before automated triage.

I

Integrity

Deterministic processing, source-health auditing and a deny-by-default publication firewall protect the public layer. Ambiguous evidence is allowed to abstain rather than being forced into a confident label.

V

Verification

Automated harvesting, validation, evidence-card infrastructure, registered-trial/publication linking, prevalence processing, research-integrity monitoring and human-review infrastructure are now implemented within the research system.

V

Validity

The system is designed for evidence discovery, surveillance, triage, structured extraction and review support. Automation is not equivalent to scientific quality appraisal, causal inference or clinical guidance.

Current project synopsis: September 2026

Achievements so far

Beyond a prototype.

The project now has enough independent infrastructure to operate as a coherent evidence system rather than a loose collection of experiments.

Research infrastructure

Mature private research repository

Dedicated sccvaping account structure, automated harvesting and validation, evidence-card and synthesis infrastructure, registered-trial/publication linking, prevalence processing and source-health auditing.

Publication control

Deny-by-default firewall

Public-safe outputs are separated from raw and investigative material, with provenance-preserving publication controls and independent Drive backup credentials.

Human accountability

Bounded ML and AI governance

ML review prioritisation, human-review infrastructure and formal AI/ML rules keep assistance distinct from scientific judgement.

Collaboration

Academic sharing structure

A curated Google Drive workspace now supports academic evidence packs, methods, exports, reports, correspondence and archive material without exposing the private engine.

Innovation lifecycle

Promote what survives evaluation, not what merely looks impressive.

A new method graduates only if it outperforms an appropriate simpler baseline without damaging provenance, reproducibility or human accountability.

DiscoverHypothesisBaselineLocked evaluationSandboxEvaluateHuman benchmarkAuditPromote / rejectMonitorRevalidate

This governance has already proved useful: an invalid experimental status was caught during the first implementation, corrected, and the subsequent main validation completed successfully.

Current innovation programme

Six concrete experiments.

AI belongs here as a tool for finding, ranking, checking, comparing and assisting — not quietly becoming the scientist.

01

Active-learning screening

Test whether intelligent prioritisation can reduce manual screening workload while maintaining recall.

02

LLM-assisted second extraction

Benchmark source-grounded extraction assistance against human-reviewed reference material.

03

Trial/publication graph ranking

Use relationships between trials, publications, DOIs, authors and institutions to improve evidence discovery and prioritisation.

04

Evidence-drift monitoring

Detect material changes, corrections, retractions and shifts in the evidence base over time.

05

Citation-grounded academic RAG

Investigate a retrieval system that answers only from traceable project evidence and preserves citation context.

06

Selective / conformal classification

Explore calibrated abstention only once sufficient independent human labels exist.

Next-stage goals

Make the infrastructure academically useful and publicly distinctive.

  • Build a substantially more interactive SCC Vaping Observatory.
  • Create a live study and evidence explorer with filters and provenance drill-down.
  • Map trial → publication → DOI → author → institution relationships.
  • Implement continuous change, correction and retraction detection.
  • Benchmark active learning against the existing TF-IDF/logistic approach.
  • Run a blinded benchmark of LLM-assisted extraction.
  • Publish downloadable reproducibility packages and model/system cards.
  • Produce a professionally presented Academic Research Pack.

Academic Research Pack

A formal, citable project package.

The planned pack will bring together the one-page synopsis, full project report, aims and research questions, methodology, evidence-map statistics, limitations, governance statement, current study register, provenance manifest, reviewer guide and citation information.

The purpose is to make external scrutiny easier: a reviewer should be able to understand what the system does, inspect how evidence reached the public layer, identify the boundaries of automation, and reproduce or challenge the relevant outputs.

Public boundary

Transparency without exposing operational secrets.

Public pages describe purpose, evidence principles, methods, validated status and intended outputs. Credentials, private working data, restricted source material, anti-abuse controls and security-sensitive implementation details remain outside the public surface.

The project does not claim that automated discovery, classification or extraction is scientific verification. Human scientific accountability remains explicit wherever conclusion-sensitive evidence is concerned.

Research standards →

SCC Vaping

Evidence that can be interrogated, verified, challenged and updated.

The project is designed to improve the speed of evidence surveillance without trading away provenance, uncertainty, reproducibility or human responsibility.

Open the SCC Vaping Observatory