Embedded systems Engineer task
Worldwide
a quick break down of what this is, and what we will need you to do: Hashlist is building a benchmark that measures how well AI models do embedded engineering work under fixed constraints. Most coding benchmarks check two things: does the code compile, and do the tests pass. Very little measures whether the code respects the constraints that actually govern embedded work, such as a fixed RAM ceiling, an interrupt latency budget, a datasheet rule about write ordering, or a bus timing deadline. Generated code breaks those constraints regularly while compiling cleanly and passing a reasonable test suite. That is the gap we are measuring, and the results are published as a report on where models hold up and where they fail. The tasks cannot be scraped or generated. They need an engineer who has made, or reviewed, the specific mistake. What a task is A task is a small folder of C that runs on a normal machine. No hardware required: if the difficulty depends on how a chip behaves, you write a mock that reproduces that behaviour. It contains: - PROMPT.md, the ticket the model is given - api.h, the function to implement - spec/, the datasheet section or standard the answer depends on, if there is one - budgets.json, the numeric limits, each recording where its number came from - tests/, a harness that runs a solution and prints what it measured - gold/, your correct solution - naive/, two wrong solutions a competent engineer would plausibly submit The last one is the part that matters most. Each wrong solution has to pass every test and still break at least one limit. That is what separates this from an ordinary test suite. What we need you to do Write two or three tasks, on problems you have actually hit in professional work. We also want to know how long each one takes you. That is a real part of what we are asking: we are sizing the work, so please keep a rough note of your time per task. An honest number is more useful to us than a fast one. You do not need to have the whole thing worked out before you start. Pick a sub-category, name the mistake, and the platform scaffolds the files for you. How it works in practice 1. You get a login. Access is by invitation, so an admin adds your address and gives you a password. 2. Press Build a task. A guide opens first, and it stays docked beside the editor while you work. You can skip it, but it is worth reading once. 3. Pick a sub-category. If your example fits one but is not among the categories listed there, use Create your own task in this subtopic and name the mistake yourself. You are not restricted to a fixed list, and slots are not exclusive, so it does not matter if someone else has taken the one you want. 4. Fill in the files. Everything saves as you go, so you can stop and come back. 5. Press Run admission. The platform compiles your solutions and runs them in a sandbox. It checks that your correct solution passes everything, and that each wrong solution passes the tests while breaking at least one limit. You cannot submit until that passes. 6. Optionally, run your task against real models to see whether it separates them. 7. Submit. Another engineer reviews it against five criteria and either accepts it or sends it back with notes. The guide inside the platform covers all of this in detail, stage by stage, with a full worked example you can open and read. The most common reasons a task comes back - The limits were chosen rather than measured. This is the most frequent one - No wrong solution actually exists, which means the tests already cover the rule - The mock is more permissive than the real component, so the wrong solution passes - The prompt names the rule that decides the answer, which reduces the task to transcription - The mistake is a typo rather than a decision Using AI You may use AI tools for parts of the work: boilerplate, harness scaffolding, tidying prose. Do not use them to invent the task. The value here is your judgement about what a real engineer gets wrong and why, and a generated task does not carry that.
$200.00
Fixed-price- ExpertExperience Level
- Remote Job
- One-time projectProject Type
Skills and Expertise
Activity on this job
- Proposals:10 to 15
- Last viewed by client:2 days ago
- Hires:6
- Interviewing:4
- Invites sent:31
- Unanswered invites:18
About the client
- FINHelsinki5:57 AM
- $290 total spent9 hires, 9 active
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by