AML- Jr. AI / ML Modelling Engineer
bnpparibas
Job Description
Direct Responsibilities
- Model Development & Evaluation – Design, train, and assess ML models (NNs, GNNs, gradient‑boosted ensembles, unsupervised anomaly detectors) for money‑laundering detection across transaction, client, and network levels.
- Feature Engineering Pipeline – Build and optimise Spark/Delta Lake pipelines, ensuring feature consistency between training and production.
- Lifecycle Management – Track experiments, version models, and run champion‑challenger workflows with MLflow.
- Low‑Latency Deployment – Deploy and maintain inference services (e.g., BentoML) on the on‑prem Kubernetes platform.
- Calibration & Scoring – Calibrate models and map scores to probabilities to deliver reliable, decision‑ready risk scores.
- Explainability & Governance – Provide model explainability for audits, document design/assumptions/validation, and guarantee full reproducibility and auditability for regulatory compliance.
- Cross‑Functional Collaboration – Partner with data engineers, graph/LLM specialists, and compliance analysts to translate detection requirements into production‑ready models.
Technical & Behavioral Competencies
- Python engineering & secure code – Expert in object‑oriented, modular, and reusable Python code; packaging, logging, metrics, and tracing (ELK, Prometheus, Grafana, OpenTelemetry); security primitives (OAuth2/JWT) and data‑privacy compliance (GDPR/CCPA).
- ML modeling & calibration – Hands‑on with PyTorch (custom architectures, focal loss, training loops, GPU optimization), XGBoost/LightGBM, scikit‑learn/Keras/TensorFlow; deep knowledge of model calibration (Platt scaling, score‑to‑probability, MC‑Dropout), evaluation (PR/AUC‑ROC, F‑beta, lift charts) and handling class imbalance along with the underlying statistical concepts (calibration theory, hypothesis testing).
- Large‑scale data processing – Proficient in Pandas, Dask, Polars (or equivalent) for transactional‑scale datasets; feature engineering, distributed pipelines, and performance‑aware data handling.
- MLOps & lifecycle management – MLflow (experiment tracking, model registry, stage promotion), BentoML (runners, adaptive batching, CPU/GPU pools), GitOps/Argo CD for A/B‑testing and champion‑challenger workflows; CI/CD pipelines with Git/Bitbucket.
- Python microservices & containerization – Designing and deploying APIs with FastAPI, Flask, or Django; async I/O, Celery task queues, Docker packaging, and Kubernetes orchestration for scalable inference services.
- Data storage & query optimization – Strong command of SQL (PostgreSQL, Oracle) and NoSQL/graph stores (Cassandra, MongoDB, Neo4j, JanusGraph); schema design, indexing, and performance tuning.
- Experience with AWS, Azure, or GCP; messaging/streaming (Kafka, RabbitMQ); cross‑team and multi‑geography collaboration, rapid technology adoption, mentorship, and end‑to‑end ownership of solutions.
- Strong collaboration and learning mindset, with the ability to work across teams and geographies, adapt to new technologies, mentor others, and take end-to-end ownership.