As a Data Scientist at Polco, you'll synthesize advanced insights for local governments, focusing on descriptive, predictive, and causal analytics to impact over 36 million Americans.
[8BE] Data Scientist (AI + ML)
📜 Description
- Design and implement Bayesian statistical models for decision-making under uncertainty.
- Build Markov chain and Hidden Markov Model formulations for sequential and behavioral patterns.
- Apply MCMC methods, including Metropolis-Hastings sampling, for estimating posterior distributions.
- Develop Gaussian Mixture Models to support segmentation use cases.
- Implement Expectation-Maximization for latent-variable estimation in mixture models.
- Collaborate with backend engineering to translate models into production service architecture.
🛠️ Requirements
- +90% English written and oral (at least B2 level) with excellent communication skills
- Strong, demonstrable background in Bayesian statistics/Bayesian inference, Markov chains, Hidden Markov Models, MCMC methods (including Metropolis-Hastings sampling), mixture models (ideally Gaussian Mixture Models), and Expectation-Maximization.
- Proven experience building and deploying statistical/ML models into production systems, not just research notebooks or offline analysis.
- Proficiency in Python (or R) with standard probabilistic/statistical libraries (e.g., PyMC, Stan, scikit-learn, NumPy/SciPy) for model development and validation.
- Ability to translate statistical/mathematical models into service-oriented production architecture — defining APIs and data contracts and working directly with backend engineers to integrate them.
- Solid understanding of version control, testing practices, and CI/CD, sufficient to collaborate effectively with an engineering team on production delivery.
- Strong written and verbal communication skills, with the ability to explain model behavior, assumptions, and uncertainty to non-technical stakeholders.
- Experience in e-commerce or retail domains, particularly pricing optimization, customer segmentation, or demand forecasting.
- Experience integrating ML models with microservices architectures (REST/GraphQL) and event-driven systems (e.g., message queues/pub-sub), and deploying to cloud infrastructure.
- Familiarity with common backend service ecosystems (e.g., .NET, Java, or Node.js) — even if modeling itself is done in Python — for a smoother handoff to the production engineering team.
✨ Benefits
- Flexible schedule and Work From Anywhere
- Referral Program
- Supportive and chill atmosphere
Full job description
Company Description
We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!
Job Description
We are seeking a Data Scientist with deep expertise in probabilistic AI and statistical machine learning to support a client's e-commerce platform, built on a distributed microservices architecture. The platform roadmap includes a set of intelligence capabilities that require rigorous statistical modeling rather than standard supervised ML. This role is responsible for designing, validating, and productionizing probabilistic models, and for working closely with backend engineering to translate those models into service-oriented production architecture within the platform's existing microservices ecosystem.
Project Length: 3 - 6 months.
What you will do
- Bayesian modeling and inference: Design and implement Bayesian statistical models — priors, likelihoods, and posterior inference — to support decisioning under uncertainty across pricing, segmentation, and demand-related use cases.
- Markov chains and Hidden Markov Models: Build Markov chain and Hidden Markov Model formulations for sequential and behavioral patterns (e.g., customer lifecycle stages, state transitions), producing outputs that downstream services can consume.
- MCMC and Metropolis-Hastings sampling: Apply Markov Chain Monte Carlo methods, including Metropolis-Hastings sampling, to estimate posterior distributions for models without closed-form solutions, and validate convergence and sampling quality.
- Mixture modeling: Develop mixture models — Gaussian Mixture Models in particular — to support segmentation use cases, identifying latent customer or product groupings from transactional and behavioral data.
- Expectation-Maximization: Implement Expectation-Maximization for latent-variable estimation underlying mixture models and related unsupervised learning tasks.
- Production translation: Work with backend engineering to translate statistical models into production service architecture — defining APIs, data contracts, and integration points within the platform's existing microservices and event-driven pipelines.
- Model lifecycle management: Define the approach for model training, validation, versioning, monitoring/drift detection, and retraining cadence once models are in production.
- Roadmap collaboration: Partner with delivery and engineering leads to size, sequence, and estimate probabilistic/statistical modeling initiatives on the product roadmap.
- Documentation and handoff: Document modeling assumptions, methodology, and validation results, and provide clear hand-off guidance so models remain maintainable by the engineering team after the engagement.
Qualifications
- +90% English written and oral (at least B2 level) with excellent communication skills
- Strong, demonstrable background in Bayesian statistics/Bayesian inference, Markov chains, Hidden Markov Models, MCMC methods (including Metropolis-Hastings sampling), mixture models (ideally Gaussian Mixture Models), and Expectation-Maximization.
- Proven experience building and deploying statistical/ML models into production systems, not just research notebooks or offline analysis.
- Proficiency in Python (or R) with standard probabilistic/statistical libraries (e.g., PyMC, Stan, scikit-learn, NumPy/SciPy) for model development and validation.
- Ability to translate statistical/mathematical models into service-oriented production architecture — defining APIs and data contracts and working directly with backend engineers to integrate them.
- Solid understanding of version control, testing practices, and CI/CD, sufficient to collaborate effectively with an engineering team on production delivery.
- Strong written and verbal communication skills, with the ability to explain model behavior, assumptions, and uncertainty to non-technical stakeholders.
Additional Information
Preferred Qualifications
- Experience in e-commerce or retail domains, particularly pricing optimization, customer segmentation, or demand forecasting.
- Experience integrating ML models with microservices architectures (REST/GraphQL) and event-driven systems (e.g., message queues/pub-sub), and deploying to cloud infrastructure.
- Familiarity with common backend service ecosystems (e.g., .NET, Java, or Node.js) — even if modeling itself is done in Python — for a smoother handoff to the production engineering team.
- Experience with MLOps tooling such as model registries, monitoring, and feature stores.
- Background in pricing science, recommendation systems, or marketing analytics.
Our Benefits
- Flexible schedule and Work From Anywhere
- Referral Program
- Supportive and chill atmosphere
We are accepting applications from LATAM countries
Similar jobs
Search more Data Scientist jobsThe Data Governance & Metadata Scientist will develop and maintain frameworks for metadata governance, ensuring compliance and enhancing data interoperability for a DoD customer.
As a Data Governance & Metadata Scientist, you will develop and sustain a scalable data ecosystem, ensuring compliance with DoD policies and enhancing data interoperability and analytics.
Financial Data Scientist
As a Financial Data Scientist, you will analyze structured and unstructured data, develop predictive models, and enhance decision-making processes for government and commercial clients.
As a Data Scientist, you will build machine learning models, analyze complex healthcare data, and implement innovative solutions to enhance model performance.
