Lead AI/ML Data Scientist- Vice president
citi
Job Description
-
Design, build, and deploy AI and machine learning models — including Agentic AI and Generative AI solutions — to solve complex reconciliation and data engineering challenges at enterprise scale.
-
Lead the end-to-end ML model development lifecycle, from requirements gathering and data preprocessing through to ensemble modeling, validation, and production integration.
-
Analyze large volumes of structured and unstructured financial data to uncover trends, patterns, and opportunities for optimization across banking platforms.
-
Define and deliver ML model roadmaps in collaboration with technical and business teams, ensuring alignment with project timelines, budgets, and Citi's architecture standards.
-
Translate complex data findings into clear visualizations and strategic recommendations that inform decisions made by senior business and technology leaders.
-
Partner with engineering, operations, and cross-functional teams to ensure seamless model integration, long-term scalability, and reliable performance in production environments.
-
Identify and communicate technology risks and their business implications, developing mitigation strategies and maintaining transparency with stakeholders at all levels.
-
Maintain comprehensive model documentation and support knowledge transfer to ensure continuity and adoption across teams.
Required Qualifications & Skills:
Technical Expertise:
-
10+ years hands-on experience in AI/ML development and big data engineering within Financial Services, Insurance, or Telecom environments
-
Expert-level proficiency in Python (scikit-learn, TensorFlow, PyTorch, Pandas, NumPy), R (caret, tidyverse, mlr3), and SQL (PostgreSQL, Oracle, MySQL)
-
Deep technical knowledge implementing supervised and unsupervised ML algorithms: linear/logistic regression, neural networks (CNN, RNN, LSTM, Transformers), k-means clustering, DBSCAN, decision trees (CART, C4.5), and ensemble methods (Random Forest, XGBoost, LightGBM, CatBoost)
-
Proven experience building and deploying Agentic AI and LLM-based solutions using:
-
LangGraph for complex agent orchestration and state management
-
LangChain for chain-of-thought reasoning and retrieval-augmented generation (RAG)
-
Agent Development Kit (ADK) for enterprise-grade autonomous agent development
-
-
Production-level experience with MLOps frameworks and infrastructure:
-
Apache Airflow for ML pipeline orchestration and workflow automation
-
Kubernetes for containerized model deployment and scaling
-
Docker for reproducible ML environments
-
-
Advanced proficiency with distributed computing technologies:
-
Apache Spark (PySpark, Spark MLlib) for large-scale data processing
-
Hadoop ecosystem (HDFS, MapReduce, YARN)
-
Apache Hive for data warehousing and SQL-on-Hadoop
-
-
Expertise with cloud-native data platforms:
-
AWS S3 for scalable data lake storage
-
Amazon Redshift for enterprise data warehousing
-
AWS SageMaker, Azure ML, or Google Vertex AI (beneficial)
-
Strong background in data reconciliation frameworks, data quality validation, and ETL/ELT pipelines for financial data processing at enterprise scale
-
-