Skip to Content
Writing test cases

Writing test cases

A test case is a set of plain-language instructions for the Reelspect testing agent. This page explains which fields the agent actually receives, what it can and cannot do on screen, and how to write instructions that give reliable results.

Test case fields

The Library and the Sandbox use the same core fields. Only Test Instructions (Prompt) (up to 5,000 characters) and text Attachments reach the agent. Title is sent only as a fallback when there are no instructions at all; Description and Tags are for people and are never sent.

A Sandbox test case also stores its Device Preset (required to launch, see Devices) and its Source:

  • URL: the page the agent opens first. Only http and https addresses are accepted, up to 2,000 characters; a missing https:// is added for you.
  • In the test instructions or an attachment: the agent reads the URL(s) from Test Instructions (Prompt) or a text attachment, not from the Description field. If no URL can be found there, the dialog warns you but still saves.
  • If no URL is saved, with either option, the agent first opens the project’s Default URL, when one is set (see Project settings).

A Library test case has no URL or device field: the run supplies them (see Runs). The exception is a test case promoted from the Sandbox, which silently keeps its Sandbox device; see Promote to Library.

What the agent receives

When a test case runs as part of a run, Reelspect builds the agent’s instructions from three layers, in this order, separated by blank lines:

  1. The test plan’s Test Instructions (Prompt), if the run belongs to a plan.
  2. The run’s Test Instructions (Prompt).
  3. The test case’s Test Instructions (Prompt).

Attachments are added in the same order: plan, then run, then test case. The agent also starts on the run’s Source URL, if one is set, or otherwise on the project’s Default URL.

In the Sandbox, the agent receives only the test case’s own instructions and attachments.

The agent does not see:

  • titles of test cases, plans and runs
  • any Description field
  • tags, comments and linked Jira issues

If the agent needs to know something, put it in the instructions.

When several test cases run in one run, the agent is also told which one it is working on (for example, test case 2 of 5).

Each execution keeps the exact instructions it received and the test case version it ran (shown as v1, v2 and so on in the execution list). Editing a test case later does not change past results.

Attachments

You can attach text files (.txt, .csv, .xml, .json, .md) to a test case, a test plan or a run, up to 20 files on each, about 3.7 MB per file. The full content of each file is added after the instructions, under the file name. Use them for longer data such as a list of game launch URLs or test accounts.

Login details and other secrets

There is no separate credentials field. Write login details into the instructions:

  • Put shared logins in the test plan’s or run’s Test Instructions (Prompt), so every test case gets them.
  • Put a login that only one test needs in that test case.

Instructions are stored as plain text and are visible to everyone with access to the project. Everything the agent types, including passwords, appears in the session’s Logs tab and in downloaded logs. Use dedicated test accounts, never real customer or personal accounts.

Requested output

You can ask the agent for structured data in addition to its written report, for example a table with one row per game. There is no separate setting: describe the shape you want in the instructions.

You ask forYou get
A table with named columns, CSV, JSON or a spreadsheetA table, one row per thing tested, columns in the order you gave
One item with named fields (a JSON object)A two-column Field / Value table
A list, a sentence or a single valueText, in the form you asked for

The result appears in the Requested Output section of a run result and of a Sandbox execution, and can be saved as CSV or JSON. The CSV opens correctly in Excel, including non-Latin text.

How to ask well:

  • Name the columns or fields, in the order you want them, for example Return a table with one row per game: | Provider | Game | Loaded | Load time (s) |.
  • Say what one row stands for. Every item tested gets a row, including the ones that failed.
  • Values are filled only from what the agent saw or measured. A value it could not measure reads not measured; something it never saw reads not observed.
  • The written report (expected and actual behaviour, failure reason, steps to reproduce) is always produced as well.

What the agent can do

The agent sees one screenshot of the screen per turn and acts on it. It does not read the page’s code, so it finds things the way a person does: by what is visible.

ActionNotes
Click or tapA single left click on desktop, a single tap on mobile.
Type and press keysTypes text into the focused field; presses keys such as Enter.
Scroll or swipeMouse-wheel scrolling on desktop, a swipe on mobile.
Open a URLMainly URLs given in your instructions.
WaitWaits for loading, animations or auto-playing features such as free spins.
Take notesRecords balances, bets and findings as it goes; notes are shown in the Session viewer.
Use timersFor time-based tests (“play for 30 minutes”) and for measuring durations.
CalculateWorks out bet steps, balance changes and similar arithmetic.
Identify the gameRecognizes the game from its splash screen and uses what Reelspect already knows about it.

Measuring durations

The agent can measure how long something takes, such as a game load, by reading a timer or clock at the start and at the end. Because it only sees the screen every few seconds, the true moment falls between two screenshots. When precision matters it reports a range, for example “between 15 and 22 seconds”. If it could not read both ends, it says so instead of estimating.

What the agent cannot do

  • Hear audio. It can check that a sound or mute button exists and changes state, not that sound plays.
  • Rotate the device between portrait and landscape, or resize the browser window.
  • Judge animations. It sees still screenshots, so it cannot assess smoothness or frame rate. It can check the state before and after, for example the symbols where the reels stopped.
  • Drag, hover, long-press, right-click or double-click. Drag-only sliders and chips, hover tooltips and context menus are out of reach. If a game also shows the information on click, the agent tries that.
  • Pinch-to-zoom or use any other multi-touch gesture.
  • Switch tabs or use browser back, forward or refresh in desktop browsers. In desktop tests, links that open a new tab or window (for example terms and conditions) cannot be followed yet. In mobile browsers the agent sees the whole browser, so it can switch tabs and use back, forward and refresh.
  • Change network conditions, such as going offline or slowing the connection.
  • Catch brief on-screen messages. Anything that appears and disappears between two screenshots, such as a short toast or a quick win banner, can be missed. Treat anything shown only briefly as unreliable, and check lasting evidence instead (balance, history, a status label).

When a test asks for one of these, the agent tests what it can and states in its report what it could not verify.

Native dialogs

  • In desktop browsers, JavaScript alert, confirm and prompt dialogs are accepted automatically. A test cannot choose Cancel in such a dialog.
  • On iOS simulators and iOS devices, when the agent taps while a native alert is open, the alert is accepted. Its text is kept in the session logs as evidence.

How the verdict is decided

The agent ends the session with a short summary of what it did and what stopped it. A separate review step then reads the whole session (every turn, the agent’s notes and the page’s network activity) and sets the verdict: Passed, Failed, Blocked or Error. See Verdicts for what each one means.

The report also mentions anything broken the agent saw that the test was not asking about.

Tips for writing instructions

  • One goal per test case. Several unrelated checks in one test make the verdict hard to read.
  • State the expected result. End with a line such as Expected: the balance decreases by the bet amount. The verdict is judged against it.
  • Name the game and provider exactly, for example “Sweet Bonanza by Pragmatic Play”. The agent uses the name to identify the game and apply what Reelspect already knows about it.
  • Put shared steps in the plan or run. Logins, the environment to use and how to reach the lobby belong in the plan’s or run’s instructions, so each test case only describes its own check.
  • Describe what is visible, such as button labels, icons and screen text, not code or selectors.
  • Be specific with numbers: the bet amount, how many spins, how long to play, which currency.
  • Define measurements. Say which event starts the clock and which stops it, for example “from tapping the game tile until the spin button is visible”.
  • Ask only for what you need. The agent does what the instructions ask and no more; it does not explore on its own.
  • Avoid the limitations above. Ask for visible evidence instead: “check the sound icon changes to muted” rather than “check the sound stops”.
  • Use a Single Session run when tests depend on each other, so a login in the first one carries over. Its order cannot be changed; see Execution strategies.
  • Try it in the Sandbox first. Launch, read the result, and use the Chat tab in the Session viewer to ask why the agent did something. Then refine and launch again.

Examples

Shared run instructions for a staging site:

Use the staging casino. Log in with username qa_player_01 and password Stg-Test-2026 before you start. Accept the cookie banner if it appears. Play in real-money mode unless the test says otherwise.

Launch the first five games per provider and check that they load:

For each of these providers: Pragmatic Play, Play'n GO, NetEnt. 1. In the lobby, filter or search for the provider. 2. Open each of the first five games listed, one at a time. 3. A game counts as loaded when the loading screen is gone and the spin button is visible. Measure the load time with a timer, from tapping the game tile until the game counts as loaded. 4. Close the game and return to the lobby before opening the next one. Expected: every game loads within 20 seconds and shows no error message. Return a table with one row per game: | Provider | Game | Loaded | Load time (s) | Error text |

Spin and check the balance:

Test Big Bass Bonanza by Pragmatic Play. 1. Wait for the game to load and close any intro screens. 2. Note the balance, then set the total bet to the lowest available value. 3. Spin once. Wait until the reels stop and any win presentation has finished. 4. Note the win amount and the new balance. Expected: new balance = old balance - bet + win, with no rounding difference. Return a JSON object with the fields startBalance, bet, win, endBalance.

Check the bet limits without spinning:

Test Gates of Olympus by Pragmatic Play in demo mode. Using the bet controls, find the minimum and the maximum total bet. Do not spin. Expected: the minimum total bet is 0.20 EUR and the maximum is 100.00 EUR. Return a table: | Limit | Expected | Observed |