Writing test cases
A test case is a set of plain-language instructions for the Reelspect testing agent. This page explains which fields the agent actually receives, what it can and cannot do on screen, and how to write instructions that give reliable results.
Test case fields
The Library and the Sandbox use the same core fields. Only Test Instructions (Prompt) (up to 5,000 characters) and
text Attachments reach the agent. Title is sent only as a fallback when there are no instructions at all;
Description and Tags are for people and are never sent.
A Sandbox test case also stores its Device Preset (required to launch, see Devices) and its
Source:
- URL: the page the agent opens first. Only
httpandhttpsaddresses are accepted, up to 2,000 characters; a missinghttps://is added for you. - In the test instructions or an attachment: the agent reads the URL(s) from
Test Instructions (Prompt)or a text attachment, not from theDescriptionfield. If no URL can be found there, the dialog warns you but still saves. - If no URL is saved, with either option, the agent first opens the project’s Default URL, when one is set (see Project settings).
A Library test case has no URL or device field: the run supplies them (see Runs). The exception is a test case promoted from the Sandbox, which silently keeps its Sandbox device; see Promote to Library.
What the agent receives
When a test case runs as part of a run, Reelspect builds the agent’s instructions from three layers, in this order, separated by blank lines:
- The test plan’s
Test Instructions (Prompt), if the run belongs to a plan. - The run’s
Test Instructions (Prompt). - The test case’s
Test Instructions (Prompt).
Attachments are added in the same order: plan, then run, then test case. The agent also starts on the run’s Source URL, if one is set, or otherwise on the project’s Default URL.
In the Sandbox, the agent receives only the test case’s own instructions and attachments.
The agent does not see:
- titles of test cases, plans and runs
- any
Descriptionfield - tags, comments and linked Jira issues
If the agent needs to know something, put it in the instructions.
When several test cases run in one run, the agent is also told which one it is working on (for example, test case 2 of 5).
Each execution keeps the exact instructions it received and the test case version it ran (shown as v1, v2
and so on in the execution list). Editing a test case later does not change past results.
Attachments
You can attach text files (.txt, .csv, .xml, .json, .md) to a test case, a test plan or a run, up to 20
files on each, about 3.7 MB per file. The full content of each file is added after the instructions, under the file
name. Use them for longer data such as a list of game launch URLs or test accounts.
Login details and other secrets
There is no separate credentials field. Write login details into the instructions:
- Put shared logins in the test plan’s or run’s
Test Instructions (Prompt), so every test case gets them. - Put a login that only one test needs in that test case.
Instructions are stored as plain text and are visible to everyone with access to the project. Everything the agent types, including passwords, appears in the session’s Logs tab and in downloaded logs. Use dedicated test accounts, never real customer or personal accounts.
Requested output
You can ask the agent for structured data in addition to its written report, for example a table with one row per game. There is no separate setting: describe the shape you want in the instructions.
| You ask for | You get |
|---|---|
| A table with named columns, CSV, JSON or a spreadsheet | A table, one row per thing tested, columns in the order you gave |
| One item with named fields (a JSON object) | A two-column Field / Value table |
| A list, a sentence or a single value | Text, in the form you asked for |
The result appears in the Requested Output section of a run result and of a Sandbox execution, and can be saved as CSV or JSON. The CSV opens correctly in Excel, including non-Latin text.
How to ask well:
- Name the columns or fields, in the order you want them, for example
Return a table with one row per game: | Provider | Game | Loaded | Load time (s) |. - Say what one row stands for. Every item tested gets a row, including the ones that failed.
- Values are filled only from what the agent saw or measured. A value it could not measure reads
not measured; something it never saw readsnot observed. - The written report (expected and actual behaviour, failure reason, steps to reproduce) is always produced as well.
What the agent can do
The agent sees one screenshot of the screen per turn and acts on it. It does not read the page’s code, so it finds things the way a person does: by what is visible.
| Action | Notes |
|---|---|
| Click or tap | A single left click on desktop, a single tap on mobile. |
| Type and press keys | Types text into the focused field; presses keys such as Enter. |
| Scroll or swipe | Mouse-wheel scrolling on desktop, a swipe on mobile. |
| Open a URL | Mainly URLs given in your instructions. |
| Wait | Waits for loading, animations or auto-playing features such as free spins. |
| Take notes | Records balances, bets and findings as it goes; notes are shown in the Session viewer. |
| Use timers | For time-based tests (“play for 30 minutes”) and for measuring durations. |
| Calculate | Works out bet steps, balance changes and similar arithmetic. |
| Identify the game | Recognizes the game from its splash screen and uses what Reelspect already knows about it. |
Measuring durations
The agent can measure how long something takes, such as a game load, by reading a timer or clock at the start and at the end. Because it only sees the screen every few seconds, the true moment falls between two screenshots. When precision matters it reports a range, for example “between 15 and 22 seconds”. If it could not read both ends, it says so instead of estimating.
What the agent cannot do
- Hear audio. It can check that a sound or mute button exists and changes state, not that sound plays.
- Rotate the device between portrait and landscape, or resize the browser window.
- Judge animations. It sees still screenshots, so it cannot assess smoothness or frame rate. It can check the state before and after, for example the symbols where the reels stopped.
- Drag, hover, long-press, right-click or double-click. Drag-only sliders and chips, hover tooltips and context menus are out of reach. If a game also shows the information on click, the agent tries that.
- Pinch-to-zoom or use any other multi-touch gesture.
- Switch tabs or use browser back, forward or refresh in desktop browsers. In desktop tests, links that open a new tab or window (for example terms and conditions) cannot be followed yet. In mobile browsers the agent sees the whole browser, so it can switch tabs and use back, forward and refresh.
- Change network conditions, such as going offline or slowing the connection.
- Catch brief on-screen messages. Anything that appears and disappears between two screenshots, such as a short toast or a quick win banner, can be missed. Treat anything shown only briefly as unreliable, and check lasting evidence instead (balance, history, a status label).
When a test asks for one of these, the agent tests what it can and states in its report what it could not verify.
Native dialogs
- In desktop browsers, JavaScript
alert,confirmandpromptdialogs are accepted automatically. A test cannot choose Cancel in such a dialog. - On iOS simulators and iOS devices, when the agent taps while a native alert is open, the alert is accepted. Its text is kept in the session logs as evidence.
How the verdict is decided
The agent ends the session with a short summary of what it did and what stopped it. A separate review step then reads the whole session (every turn, the agent’s notes and the page’s network activity) and sets the verdict: Passed, Failed, Blocked or Error. See Verdicts for what each one means.
The report also mentions anything broken the agent saw that the test was not asking about.
Tips for writing instructions
- One goal per test case. Several unrelated checks in one test make the verdict hard to read.
- State the expected result. End with a line such as
Expected: the balance decreases by the bet amount.The verdict is judged against it. - Name the game and provider exactly, for example “Sweet Bonanza by Pragmatic Play”. The agent uses the name to identify the game and apply what Reelspect already knows about it.
- Put shared steps in the plan or run. Logins, the environment to use and how to reach the lobby belong in the plan’s or run’s instructions, so each test case only describes its own check.
- Describe what is visible, such as button labels, icons and screen text, not code or selectors.
- Be specific with numbers: the bet amount, how many spins, how long to play, which currency.
- Define measurements. Say which event starts the clock and which stops it, for example “from tapping the game tile until the spin button is visible”.
- Ask only for what you need. The agent does what the instructions ask and no more; it does not explore on its own.
- Avoid the limitations above. Ask for visible evidence instead: “check the sound icon changes to muted” rather than “check the sound stops”.
- Use a Single Session run when tests depend on each other, so a login in the first one carries over. Its order cannot be changed; see Execution strategies.
- Try it in the Sandbox first. Launch, read the result, and use the Chat tab in the Session viewer to ask why the agent did something. Then refine and launch again.
Examples
Shared run instructions for a staging site:
Use the staging casino. Log in with username qa_player_01 and password Stg-Test-2026 before you start.
Accept the cookie banner if it appears. Play in real-money mode unless the test says otherwise.Launch the first five games per provider and check that they load:
For each of these providers: Pragmatic Play, Play'n GO, NetEnt.
1. In the lobby, filter or search for the provider.
2. Open each of the first five games listed, one at a time.
3. A game counts as loaded when the loading screen is gone and the spin button is visible.
Measure the load time with a timer, from tapping the game tile until the game counts as loaded.
4. Close the game and return to the lobby before opening the next one.
Expected: every game loads within 20 seconds and shows no error message.
Return a table with one row per game: | Provider | Game | Loaded | Load time (s) | Error text |Spin and check the balance:
Test Big Bass Bonanza by Pragmatic Play.
1. Wait for the game to load and close any intro screens.
2. Note the balance, then set the total bet to the lowest available value.
3. Spin once. Wait until the reels stop and any win presentation has finished.
4. Note the win amount and the new balance.
Expected: new balance = old balance - bet + win, with no rounding difference.
Return a JSON object with the fields startBalance, bet, win, endBalance.Check the bet limits without spinning:
Test Gates of Olympus by Pragmatic Play in demo mode.
Using the bet controls, find the minimum and the maximum total bet. Do not spin.
Expected: the minimum total bet is 0.20 EUR and the maximum is 100.00 EUR.
Return a table: | Limit | Expected | Observed |