Scikit-Learn compatible transformer that turns categorical variables into dense entity embeddings.
-
Updated
Aug 14, 2023 - Jupyter Notebook
Scikit-Learn compatible transformer that turns categorical variables into dense entity embeddings.
Lightning-fast data preprocessing and feature engineering for Python, built on Polars. 108 sklearn-style transformers for imputation, encoding, scaling, outlier clipping, discretization and feature selection, with Pipeline support and ONNX export for low-latency inference.
End-to-End Feature Engineering for Machine Learning
Automatic optimal discretization pipeline
Open source machine learning library with various machine learning tools
Weight of Evidence Encoding & Information Value
Perform semi automated exploratory data analysis, feature engineering and feature selection on provided dataset by visualizing every possibilities on each step and assisting the user to make a meaningful decision to achieve a low-bias and low-variance model.
[ECAI' 25]: MOCA-HESP: Meta High-dimensional Bayesian Optimization for Combinatorial and Mixed Spaces via Hyper-ellipsoid Partitioning
A novel deep learning architecture for tabular data that treats categorical features in a discrete quantized space, inspired by Vector Quantization from VQ-VAE.
This repository is a comprehensive guide to different Encoding techniques in Machine Learning, explaining when to use each method and best practices. You'll find practical examples, ready-to-use code, and comparisons between various techniques like Label Encoding, One-Hot Encoding, Target Encoding, and more!
A machine learning project to predict loan defaults in a German bank's customer base. Using the German Credit Risk dataset, it explores key factors contributing to defaults and trains models like Random Forest, GBM, and XGBoost. Includes EDA, data processing, hyperparameter tuning, and model evaluation.
House Price Prediction is a machine learning workflow for estimating housing prices, featuring data cleaning, EDA, feature scaling, categorical encoding, and multiple regression models (Linear, Ridge, Decision Tree, Gradient Boosting, XGBoost, CatBoost, SVR) implemented in Python with Scikit-learn.
Binary classifier with Venn-ABERS calibration, temporal validation, hyper-parameter optimisation, and fraud detection out of the box
This was a challenge that predicts a marketplace promotion using a recommendation system prediction.
Housing Prices Prediction using Machine Learning Developed a regression model to predict housing prices using data preprocessing, feature engineering, and various regression algorithms. Tuned hyperparameters and evaluated performance with key metrics (RMSE, MAE, R²).
Stacked Classifier
Upstream feature_engine v1.9.4 plus one added OOFMeanEncoder (leak-free out-of-fold target encoding), proposed back to feature-engine/feature_engine as issue #1050. An upstream snapshot, not a standalone project.
A production-grade Machine Learning application predicting startup valuations based on Shark Tank India Seasons 1–3 data. Built with Multiple Linear Regression (OLS), the project features interactive dashboards, benchmark analytics, and a health radar for startup metrics.
Machine Learning Models
To associate your repository with the categorical-encoding topic, visit your repo's landing page and select "manage topics."