Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

EntropyGameTheoryEngine 🧠🤖

Universal Probability Audit, Entropy Space Decomposition & Autonomous Data Quality Auditing Engine

Python ML Status


🌐 Choose Your Language / Выберите язык / Seleccione el idioma


🇬🇧 English: Core Concept & Discovery

This project introduces an invariant tabular machine learning methodology based on game theory and the decomposition of noisy spaces. During extreme stress testing (injecting 20% random noise into the target variable), a unique indirect benefit was discovered: the framework acts as an autonomous Data Quality Auditor.

Mathematical Metric of Data Noise:

When the Monolithic Triad of gradient boostings is trained on clean geometric invariants (Ratios), the gap between the surface and deep AUC scores becomes a powerful diagnostic signal:

  • Scores converging near ~1.0: Sterile dataset, perfect labeling quality.
  • Surface AUC dropping to ~0.70 while Deep AUC remains >0.97: A direct mathematical indicator that the target variable contains roughly 20% random errors (dirty labels). The engine hits the theoretical entropy ceiling, flagging labeling anomalies for targeted human review (Active Learning).

🛠 Modular System Architecture

  1. Bayesian Mapper: Identifies regions of "pure bluffing" in distributions (probability entropy hovering between $0.40 - 0.60$).
  2. Invariant Extractor: Transforms absolute noisy metrics into relative non-linear synergy coefficients, erasing localized noise.
  3. Monolithic Triad (LGBM + XGB + CatBoost): A high-capacity ensemble designed for maximum generalization power.
  4. Rank Post-Processing: Maps raw probabilities into rank space using Spearman's method, ensuring total immunity to scale shifts on private test sets.

🇷🇺 Русский: Концепция и Открытие

Протокол аудита разметки и устойчивость к Private Shake-up

Проект предлагает инвариантную методологию табличного машинного обучения, основанную на теории игр и декомпозиции зашумленных пространств. В ходе стресс-тестирования системы с экстремальным внедрением 20% случайного шума в целевую переменную было обнаружено уникальное косвенное преимущество фреймворка: он способен выступать в роли автономного детектора качества разметки данных (Data Quality Auditor).

Математический маркер зашумленности:

Когда Триада тяжелых бустингов получает на вход чистые геометрические инварианты (Ratios), разница между поверхностным и глубинным скором становится диагностическим сигналом:

  • Сближение скоров к ~1.0: Стерильный датасет, идеальное качество разметки.
  • Падение поверхностного AUC к ~0.70 при сохранении глубинного AUC >0.97: Прямой математический индикатор того, что в целевом признаке присутствует около 20% случайных ошибок (грязных меток). Движок упирается в теоретический потолок энтропии и подсвечивает аномалии разметки для точечного аудита (Active Learning).

🛠 Модульная Архитектура Системы

  1. Байесовский картограф: Находит зоны «чистого блефа» распределения (энтропия вероятностей в районе $0.40 - 0.60$).
  2. Экстрактор инвариантов: Переводит абсолютные зашумленные метрики в относительные нелинейные коэффициенты синергии, стирая локальный шум.
  3. Монолитная Триада (LGBM + XGB + CatBoost): Ансамбль, обеспечивающий максимальную обобщающую способность.
  4. Ранговый пост-процессинг: Маппинг вероятностей в пространство рангов через метод Спирмена для полной защиты от сдвига масштаба на приватных тестах.

🇪🇸 Español: Concepto Central y Descubrimiento

Inmunidad al Private Shake-up y Auditoría de Etiquetas

Este proyecto introduce una metodología invariante de aprendizaje automático tabular basada en la teoría de juegos y la descomposición de espacios ruidosos. Durante las pruebas de estrés extremo (inyección del 20% de ruido aleatorio en la variable objetivo), se descubrió una ventaja indirecta única: el marco funciona como un Auditor Autónomo de Calidad de Datos (Data Quality Auditor).

Marcador Matemático de Ruido en los Datos:

Cuando el Triplete Monolítico de boostings se entrena con invariantes geométricos puros (Ratios), la brecha entre el AUC superficial y el AUC profundo se transforma en una señal de diagnóstico:

  • Convergencia de puntuaciones cerca de ~1.0: Conjunto de datos estéril, calidad de etiquetado perfecta.
  • Caída del AUC superficial a ~0.70 mientras que el AUC profundo se mantiene >0.97: Un indicador matemático directo de que la variable objetivo contiene aproximadamente un 20% de errores aleatorios (etiquetas sucias). El motor choca contra el techo teórico de entropía, aislando las anomalías para una auditoría humana específica (Active Learning).

🛠 Arquitectura Modular del Sistema

  1. Mapeador Bayesiano: Identifica regiones de "puro farol" en las distribuciones (entropía de probabilidad entre $0.40 - 0.60$).
  2. Extractor de Invariantes: Transforma métricas ruidosas absolutas en coeficientes de sinergia no lineales relativos, eliminando el ruido local.
  3. Triplete Monolítico (LGBM + XGB + CatBoost): Un ensamble de alta capacidad diseñado para la máxima potencia de generalización.
  4. Postprocesamiento de Rangos: Mapea las probabilidades brutas en el espacio de rangos mediante el método de Spearman, garantizando inmunidad total frente a cambios de escala en pruebas privadas.

machine-learning, tabular-data, game-theory, gradient-boosting, catboost, xgboost, lightgbm, data-quality, data-auditing, anomaly-detection, credit-scoring, fraud-detection, kaggle, robust-ml, ensemble-learning, python, jupyter-notebook

About

Universal framework for tabular ML based on entropy decomposition, game theory, and autonomous data quality auditing

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages