Slim Frikha
I'm a Senior Machine Learning Research & Applied Engineer with 12+ years of experience building and scaling production-grade AI systems across early-stage startups, billion-dollar scale-ups, and national AI initiatives.
I develop models and build the training, evaluation, and deployment systems around them, making results reproducible and fast to iterate on. My experience spans NLP, deep learning, and LLMs across product-facing teams and research labs, including leading ML teams and setting research direction. Granted US patents and named author on published LLM technical reports.
Open-source agentic RAG over a personal arXiv library. ReAct agent over a three-stage retriever (dense recall, cross-encoder reranking, adaptive score-curve cutoff) with hybrid dense/BM25 fusion and swappable embedder, reranker, and LLM backends. Includes an evaluation harness that generates span-anchored QA sets and sweeps chunking and retrieval parameters to emit a tuned config.
A PyTorch native platform for training generative AI models.
Volcano Engine Reinforcement Learning for LLMs.
A flexible and efficient training framework for large-scale alignment tasks.
A framework for few-shot evaluation of language models.
Automatic evals for LLMs.
Senior ML Research Engineer · Technology Innovation Institute
Own the training, evaluation, and data tooling behind the Falcon LLM family (Falcon 3, Falcon-H1, Falcon-H1R), working at the layer between ML research, software engineering, and platform to turn one-off training, evaluation, and data work into documented, reusable workflows that researchers run themselves.
- Training infrastructure & recipes: Designed and maintained end-to-end distributed training recipes across Megatron-LM, Torchtitan, and veRL, supporting pre-training runs on up to 1,024 H100 GPUs across 14T tokens; built reproducibility workflows covering artifact tracking (configs, checkpoints, logs) and infrastructure-as-code templates (Docker, Kubernetes).
- Evaluation platform: Architected and deployed an automated evaluation toolkit featuring checkpoint-triggered runs, SQL-based result storage, Metabase dashboards, and internal model leaderboards, cutting evaluation turnaround from days to hours across three model-training teams and eliminating a manual process that had previously consumed dedicated engineer time.
- Training data tooling: Built a streaming cleaner for chat-format mid-training and post-training data, normalizing heterogeneous sources into one canonical schema and validating structure, quality, and tool-call consistency, with resumable shard-parallel runs and per-row drop-reason reporting; adopted by the 5-person data team as the standard ingestion path across 40+ datasets, removing up to 10% corrupted rows.
- Model lifecycle support: Supported pre-training, mid-training, long-context extension, and post-training, alongside inference and evaluation, across model families spanning 0.5B–34B parameters and 30+ released checkpoints; named author on the Falcon-H1 and Falcon-H1R technical reports.
- Applied AI research: Delivered a customer-facing fine-tuning initiative enhancing tool-calling in Falcon 3 10B, from requirements gathering and solution design through development and deployment.
- Open-source contributions: Upstreamed targeted improvements to Torchtitan, ChatLearn, veRL, lm-evaluation-harness, and evalchemy.
Senior ML Applied Engineer / Science & Tech Lead · Contentsquare
Led a team of 4 ML engineers advancing semantic web understanding with NLP and LLMs, cutting enterprise client onboarding from weeks to days by replacing manual setup with automated content structuring, with projected savings of $1M in 2024.
- Designed and deployed deep learning models for large-scale web content structuring, boosting data enrichment and product features.
- Oversaw ML projects across hiring, mentoring, setting research direction, aligning with product goals, solution development, deployment and improvement based on user feedback.
- Partnered with product, engineering, and design outside the ML team to move research from prototype into scalable production systems, targeting onboarding latency as a driver of time-to-value and customer retention.
- Contributed to patent filings and guided research strategy through scientific literature reviews.
Senior ML Applied Engineer · Contentsquare
Contributed to the core R&D AI team, building deep learning solutions to extract and structure web content at scale for predictive UX analytics.
- Led projects on DOM tree classification and unsupervised URL clustering for semantic layout analysis, raising average classification accuracy by 20+ points over an outsourced third-party solution and bringing the capability in-house, retiring a $200K vendor contract.
- Managed ML pipelines from product requirements to production deployment.
Lead ML Engineer · HrFlow.ai (formerly Riminder)
Led ML efforts for intelligent recruiting tools using deep learning in NLP and computer vision.
- Developed and deployed core parsing models covering layout analysis, named-entity recognition (NER), and semantic extraction, turning unstructured CVs into structured candidate profiles, enabling structure-aware matching and ranking.
- Collaborated on AI roadmap and SaaS integration in a small early-stage team.
Data Scientist · EY
Worked on data-driven consulting projects across retail and finance sectors, delivering insights through machine learning, data engineering, and visualization.
Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
How Contentsquare reduced TensorFlow inference latency with TensorFlow Serving on Amazon SageMaker
ENSTA Paris, Institut Polytechnique de Paris
Top-5 French engineering school. Coursework in information-systems architecture, software architecture, and security. Engineering double degree.