Modulai logo
Posted 1w agoStockholm, Sweden, SE | Göteborg, Sweden, SE

Master Thesis Project - 2027

MiddleRemote~€6,400 – €8,400 / mo
Required Skills
PythonC++
Job Description

1. Bridging the sim-to-real Gap: Domain randomization and synthetic data for robotic arm control (STHLM)

Background & Description

We offer a master's thesis project on closing the gap between simulation and physical hardware in robot learning. Policies trained purely on simulated data, including Vision-Language-Action (VLA) models, often fail on real hardware due to mismatches in visuals, physics, and sensing. Domain randomization addresses this by varying simulation parameters during data generation so the model learns features invariant to the sim-real gap rather than simulator artifacts (Tobin et al., 2017).

This thesis treats synthetic data production as an optimization problem: which parameter distributions, quantities, and sim/real mixtures maximize real-world performance per unit of data and compute? Recent work shows that even simple sim/real co-training recipes substantially improve manipulation success rates (Maddukuri et al., 2025), but principled, mathematically grounded strategies remain an open question.

Core idea: synthetic data is generated via a mathematically justified and data-efficient randomization strategy, a VLA or comparable model is trained on it, and the resulting policy is deployed on physical robotic arm hardware. The theoretical contribution is the mathematical analysis behind the strategy, for example coverage guarantees, sample complexity, or framing parameter selection as an optimization problem. The applied contribution is validating the policy on real hardware.

Students will work with a physical robotic arm, GPU compute, and guidance from Modulai's ML engineers

Example directions
  • Formal analysis of how domain randomization ranges should be chosen relative to the true (unknown) distribution of real-world conditions, and what guarantees this gives on real-world generalization

  • Optimizing the mixture and scheduling of simulated versus real demonstration data during VLA fine-tuning

  • Automatic or learned domain randomization, where randomization parameters are adapted based on validation performance rather than fixed by hand

  • End-to-end evaluation: train in simulation only, train with a randomization strategy, and train with sim+real co-training, then compare real-arm task success rates

ML Techniques and Tools
  • Python, PyTorch, Git, Hugging Face

  • Robotics simulators (e.g. MuJoCo, Isaac Sim, or similar) for synthetic data generation

  • Domain randomization and sim-to-real transfer methods

  • Vision-Language-Action models and other end-to-end control architectures

  • Statistical and optimization methods for data generation strategy design

  • Real-time control and deployment on physical robotic arm hardware

References

Tobin et al., Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World, 2017. arXiv:1703.06907 - https://arxiv.org/abs/1703.06907


Maddukuri et al., Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation, 2025. arXiv:2503.24361 - https://arxiv.org/abs/2503.24361

2. Image representations in property valuation (Computer vision/Tabular) (Sthlm/Gbg)

Background & Description

We offer a master's thesis project on using image data to improve the accuracy of automated property valuation. The project is run together with a growing startup that is building a state-of-the-art valuation engine, with guidance from Modulai's ML engineers.

Automated valuation models traditionally rely on tabular data: living area, number of rooms, location, construction year and historical transactions. Two apartments with near-identical records can still differ substantially in market value because of condition, renovation standard, light, layout and view. Much of that residual signal is present in listing photographs, floor plans and aerial imagery, but is rarely exploited beyond coarse heuristics. Early work showed that a learned "luxury level" derived from interior and exterior photos, combined with metadata, can outperform established metadata-only estimates (Poursaeed et al., 2017), and later studies confirm that visual features add predictive power on top of strong tabular baselines (Kostic & Jevremović, 2021).

This thesis treats the image side as a representation and integration problem: which visual representations carry the signal that tabular features miss, and how should they be fused into a production valuation model without hurting robustness, calibration or explainability? The candidate representations span the full toolbox - image classification (room type, condition, renovation standard), semantic segmentation (materials, surfaces, greenery, floor-plan geometry), object detection (fireplaces, appliances, balconies) and general-purpose embeddings from pretrained vision or vision-language backbones.

Core idea: Explore image model approaches to represent the image information as efficiently as possible, while keeping explainability of the model. Investigate how such methods may contribute to improved valuation accuracy for apartments and house listings, The methodological contribution is the comparison and fusion strategy; the applied contribution is a validated improvement in a system that is actually shipped.

Students will work with large-scale real listing data, GPU compute, and close guidance from both the company's ML team and Modulai's ML engineers. It is also a domain that is unusually easy to relate to - everyone lives somewhere.

Example directions
  • Systematic comparison of representation families - classification heads, segmentation masks, detected objects and raw embeddings - on equal footing, measured as marginal accuracy over a strong tabular baseline

  • Off-the-shelf vision-language models used as attribute extractors versus purpose-trained models and learned embeddings: accuracy, cost and latency per valuation

  • Fusion architecture: late fusion of pooled embeddings into a gradient-boosted model versus end-to-end multimodal training, including how to aggregate a variable number of images per property

  • Weak supervision from price residuals: learning image representations directly against the part of the price the tabular model cannot explain, instead of relying on generic pretrained features

  • Floor plans as structured input: extracting layout, room adjacency and geometry, and testing whether structure beats appearance

  • Robustness to presentation: staged photos, wide-angle lenses, HDR and photographer differences - separating property quality from marketing quality

  • Uncertainty and explainability: does image data mainly shift the point estimate or tighten prediction intervals, and can per-image contributions be attributed in a way a human valuer would accept?

ML Techniques and Tools
  • Python, PyTorch, Git, Hugging Face

  • CNN and vision-transformer backbones; pretrained embeddings (e.g. CLIP, DINOv2) and vision-language models

  • Semantic segmentation and object detection (e.g. SAM, Mask R-CNN, YOLO-family models)

  • Multimodal fusion and gradient boosting (LightGBM/XGBoost) over combined tabular and image features

  • Explainability (SHAP, attention and saliency maps) and uncertainty quantification (quantile regression, conformal prediction)

  • Cloud GPU compute, experiment tracking and evaluation on real data

References

Poursaeed et al., Vision-based Real Estate Price Estimation, 2017. arXiv:1707.05489 - https://arxiv.org/abs/1707.05489

Kostic & Jevremović, What Image Features Boost Housing Market Predictions?, 2021. arXiv:2107.07148 - https://arxiv.org/abs/2107.07148

Zillow neural network estimate
https://www.zillow.com/news/building-the-neural-zestimate/

3. Open Application within Applied Machine Learning

Applied Machine Learning projects encompass a wide range of domains, including healthcare, finance, natural language processing, computer vision, and more. This open application invites students to choose projects aligned with their interests and career goals. Do you have an idea - let us know what it's about by describing it. 

Required Skills

Finishing a master's in machine learning or a master's in another field but with courses in machine learning and programming added

Please include the following in your application:

  • Link to relevant GitHub account if available.

  • Grades for bachelor's and master's.

  • Updated CV or an updated LinkedIn profile.

*Suitable candidates will be called to one interview before making a final decision.
The last date for application will be the
31th of October, but if suitable candidates apply, the process will end beforehand.

About Modulai

Modulai’s clients range from startups to multinational companies. They all share that machine learning is central to how they operate, compete, and create value.

Our services range from advisory projects and feasibility studies to end-to-end development and refinement of machine learning systems and products.

We use state-of-the-art techniques, always focusing on maximizing business impact, delivering solutions in areas such as credit risk, fraud detection, dynamic pricing, recommendation systems, computer vision, natural language processing, opportunity spotting, logistics optimization, up-sell, cross-sales, smart building optimization, predictive maintenance, and route planning.

Other 

When doing a master thesis project at Modulai, you are invited to all team activities such as daily stand-ups, weekly learning breakfasts, monthly AWs, and other team activities. We look forward to having you as part of our team! 

Ready to apply? Optimize your CV for this specific jobAI customizes your experience bullets and increases chances to get hired.

Similar Openings in AI & Machine Learning

View all in category