Feature Selection for Bioinformatics with Python
01 September 2026
07 September 2026
For-profit: 500 CHF
Overview
In bioinformatics and computational biology, researchers routinely work with high-dimensional datasets where the number of features (genes, proteins, variants, metabolites) far exceeds the number of samples. This "large p, small n" scenario presents unique challenges for building useful models and extracting biological insights from the data. Feature selection methods help identify the most informative variables, helping build models with better generalisability, better interpretability, and revealing biomarkers.
This 1-day course provides a hands-on introduction to feature selection methods specifically tailored for bioinformatics applications. Participants will learn both classical and state-of-the-art approaches with a focus on their practical application to real biological datasets.
The course will also cover when and why to use different approaches, how to validate selected features, and common pitfalls in high-dimensional biological data analysis.
Audience
This course is designed for PhD students, postdoctoral and other researchers in the life sciences from both academia and industry who work with high-dimensional biological data and want to improve their predictive modelling and biomarker discovery workflows.
Learning Objectives
At the end of the course, participants will be able to:
- Explain the fundamental challenges of feature selection in high-dimensional biological data (curse of dimensionality, spurious correlations, multiple testing)
- Distinguish between filter, wrapper, and embedded feature selection methods and select appropriate approaches for different biological questions
- Implement and apply basic and advanced feature selection methods using scikit-learn and complementary Python packages (VarianceThreshold, L1 penalty, Recursive Feature Elimination, Boruta, knockoffs framework)
- Recognise and avoid common pitfalls in feature selection pipelines, such as data leakage and multicollinearity
- Implement stability selection to assess feature selection robustness.
Prerequisites
Knowledge/competencies:
The course is intended for people already familiar with Python programming (including NumPy and Pandas) and familiar with different omics data. Participants should be comfortable with fundamental machine learning concepts: supervised learning (classification/regression), train/test splits, cross-validation, overfitting, and regularization (L1/L2 penalties). Basic statistical knowledge (hypothesis testing, p-values, multiple testing correction, FDR) is expected.
This course is part of the Machine Learning learning path. To get the most out of this course, you should meet the learning outcomes of the First Steps with Python in Life Sciences and Introduction to Machine Learning with Python courses. Upon completion of this course, you may wish to attend Ensuring More Accurate, Generalisable, and Interpretable Machine Learning Models for Bioinformatics, Diving into Deep Learning - Theory and Applications with PyTorch and Federated Learning in Bioinformatics courses.
Technical:
Your laptop must have a recent Python version (minimum 3.10) and several Python libraries installed. The needed libraries will be indicated in the course GitHub repo in due time.
Schedule – CE(S)T time zone
| Time | Module |
|---|---|
| 09:00-09:15 | Welcome & Foundations: Why feature selection matters in bioinformatics |
| 09:15-09:45 | Filters: variance threshold, correlation, mutual information |
| 09:45-10:30 | Wrapper methods: Recursive Feature Elimination (RFE) and Sequential Feature Selection (SFS) |
| 10:45-12:00 | Embedded methods: L1 regularisation, tree-based importance |
| 13:00-13:45 | Boruta : all-relevant feature selection |
| 13:45-14:30 | Knock-offs: tangling with the feature dependency problem |
| 14:45-15:45 | Practical considerations: multicollinearity, data leakage, stability selection, SHAP |
| 16:00-17:00 | Small project |
Application
The registration fees for academics are 100 CHF and 500 CHF for for-profit companies.
While participants are registered on a first come, first served basis, exceptions may be made to ensure diversity and equity, which may increase the time before your registration is confirmed.
Applications will close on 01/09/2026 or as soon as the places will be filled up. Cancellation after 07/09/2026 will not be reimbursed. Please note that participation in SIB courses is subject to our general conditions.
You will be informed by email of your registration confirmation. Upon reception of the confirmation email, participants will be asked to confirm attendance by paying the fees within 5 working days.
Venue and Time
This course will be streamed.
The course will start at 9:00 CET and end around 17:00 CET.
Precise information will be provided to the registered participants in due time.
Additional information
Coordination: Patricia Palagi, SIB Training group.
A Certificate of Attendance will be sent provided you were present at the course, whereas a Certificate of Achievement recommending 0.25 ECTS will be sent provided you passed the exam.
You are welcome to register to the SIB courses mailing list to be informed of all future courses and workshops, as well as all important deadlines using the form here.
SIB abides by the ELIXIR Code of Conduct. Participants of SIB courses are also required to abide by the same code.
For more information, please contact training@sib.swiss.