Define one or more data blocks by either start and end time or start and end scan number.
The blocks can be provided as vectors in the individual parameters or as a blocks_table - whichever is more convenient.
If you want to make segments in the blocks (optional), note that this function - manually defining blocks - removes all block segmentation. Make sure to call orbi_segment_blocks() only after finishing block definitions.
Usage
orbi_define_blocks(
dataset,
start_time.min = NULL,
end_time.min = NULL,
start_scan.no = NULL,
end_scan.no = NULL,
block_name = NA_character_,
in_filename = NULL,
blocks_table = NULL
)Arguments
- dataset
An aggregated dataset or a data frame of peaks (i.e. works directly after
orbi_identify_isotopocules()as well as with a tibble from orbi_get_data(peaks = everything()) or when reading from an IsoX file)- start_time.min
start time of the block(s), a single value or a vector for multiple blocks
- end_time.min
end time of the block(s), a single value or a vector for multiple blocks
- start_scan.no
start scan of the block(s), a single value or a vector for multiple blocks
- end_scan.no
end scan of the block(s), a single value or a vector for multiple blocks
- block_name
name(s) for the block(s), a single value or a vector for multiple blocks,
NAby default (i.e. unnamed)- in_filename
the file(s) to add the block(s) to, a single value or a vector for multiple blocks,
NA(the default if not provided) adds a block to all files. This must be thefilenameof the file(s) as it appears in thedataset.- blocks_table
alternative to the individual parameters: a data frame with any of the columns
start_time.min,end_time.min,start_scan.no,end_scan.no,block_nameandin_filename, one row per block. Any other columns are ignored. If provided, the individual parameters are not used.
Value
A data frame (tibble) with the block definitions added. Any data that is not part of a block will be marked with the value of orbi_get_option("data_type_unused"). Any previously applied segmentation will be discarded (segment column set to NA) to avoid unintended side effects.
Details
Each block has to be defined either by time (start_time.min and end_time.min) or by scan number (start_scan.no and end_scan.no) but not by both. Different blocks in the same call can use different definitions.
By default, each block is added to all files in the dataset. To add a block only to a specific file, provide its in_filename. Since all parameters are recycled, the same block definition can be added to several (but not all) files by providing a vector of filenames, e.g. in_filename = c("file1", "file2") defines the block once for each of these two files.
Blocks are matched against the scans of each file separately. A block whose range reaches beyond what a file recorded is trimmed to the scans that exist in that file, and a block that does not overlap a file at all is not added there - this is reported with a warning naming the affected files and the range each of them covers. The summary message lists the scans and times each block actually ended up covering, which is what to check if a block was trimmed.
Examples
fpath <- system.file("extdata", "testfile_flow.isox", package = "isoorbi")
df <- orbi_read_isox(file = fpath) |> orbi_simplify_isox()
#> ✔ [14ms] orbi_read_isox() loaded 6449 peaks for 1 compound (HSO4-) with 5
#> isotopocules (M0, 33S, 17O, 34S, and 18O) from testfile_flow.isox
#> ✔ [4ms] orbi_simplify_isox() kept columns filepath, filename, scan.no,
#> time.min, compound, isotopocule, ions.incremental, tic, and it.ms
# a single block
df |> orbi_define_blocks(start_time.min = 0.2, end_time.min = 0.8)
#> ✔ [41ms] orbi_define_blocks() added 1 block to 3 files
#> → block: covers scans 87 to 344 (0.202 to 0.799 min) in 3 files
#> # A tibble: 6,449 × 14
#> filepath filename scan.no time.min compound isotopocule ions.incremental
#> <chr> <fct> <int> <dbl> <fct> <fct> <dbl>
#> 1 /home/runner… s3744 1 0.002 HSO4- M0 37803.
#> 2 /home/runner… s3744 1 0.002 HSO4- 33S 326.
#> 3 /home/runner… s3744 1 0.002 HSO4- 17O 61.9
#> 4 /home/runner… s3744 1 0.002 HSO4- 34S 2215.
#> 5 /home/runner… s3744 1 0.002 HSO4- 18O 436.
#> 6 /home/runner… s3744 2 0.004 HSO4- M0 42346.
#> 7 /home/runner… s3744 2 0.004 HSO4- 33S 339.
#> 8 /home/runner… s3744 2 0.004 HSO4- 17O 66.6
#> 9 /home/runner… s3744 2 0.004 HSO4- 34S 2315.
#> 10 /home/runner… s3744 2 0.004 HSO4- 18O 438.
#> # ℹ 6,439 more rows
#> # ℹ 7 more variables: tic <dbl>, it.ms <dbl>, data_group <int>, block <int>,
#> # block_name <chr>, data_type <chr>, segment <int>
# several blocks at once
df |> orbi_define_blocks(
start_time.min = c(0.1, 0.5),
end_time.min = c(0.4, 0.8),
block_name = c("first", "second")
)
#> ✔ [67ms] orbi_define_blocks() added 2 blocks to 3 files
#> → block first: covers scans 43 to 172 (0.1 to 0.399 min) in 3 files
#> → block second: covers scans 216 to 344 (0.502 to 0.799 min) in 3 files
#> # A tibble: 6,449 × 14
#> filepath filename scan.no time.min compound isotopocule ions.incremental
#> <chr> <fct> <int> <dbl> <fct> <fct> <dbl>
#> 1 /home/runner… s3744 1 0.002 HSO4- M0 37803.
#> 2 /home/runner… s3744 1 0.002 HSO4- 33S 326.
#> 3 /home/runner… s3744 1 0.002 HSO4- 17O 61.9
#> 4 /home/runner… s3744 1 0.002 HSO4- 34S 2215.
#> 5 /home/runner… s3744 1 0.002 HSO4- 18O 436.
#> 6 /home/runner… s3744 2 0.004 HSO4- M0 42346.
#> 7 /home/runner… s3744 2 0.004 HSO4- 33S 339.
#> 8 /home/runner… s3744 2 0.004 HSO4- 17O 66.6
#> 9 /home/runner… s3744 2 0.004 HSO4- 34S 2315.
#> 10 /home/runner… s3744 2 0.004 HSO4- 18O 438.
#> # ℹ 6,439 more rows
#> # ℹ 7 more variables: tic <dbl>, it.ms <dbl>, data_group <int>, block <int>,
#> # block_name <chr>, data_type <chr>, segment <int>
# the same via a blocks table, here mixing time and scan definitions
df |> orbi_define_blocks(
blocks_table = tibble::tibble(
start_time.min = c(0.1, NA),
end_time.min = c(0.4, NA),
start_scan.no = c(NA, 200),
end_scan.no = c(NA, 350),
block_name = c("first", "second")
)
)
#> ✔ [61ms] orbi_define_blocks() added 2 blocks to 3 files
#> → block first: covers scans 43 to 172 (0.1 to 0.399 min) in 3 files
#> → block second: covers scans 200 to 350 (0.464 to 0.813 min) in 3 files
#> # A tibble: 6,449 × 14
#> filepath filename scan.no time.min compound isotopocule ions.incremental
#> <chr> <fct> <int> <dbl> <fct> <fct> <dbl>
#> 1 /home/runner… s3744 1 0.002 HSO4- M0 37803.
#> 2 /home/runner… s3744 1 0.002 HSO4- 33S 326.
#> 3 /home/runner… s3744 1 0.002 HSO4- 17O 61.9
#> 4 /home/runner… s3744 1 0.002 HSO4- 34S 2215.
#> 5 /home/runner… s3744 1 0.002 HSO4- 18O 436.
#> 6 /home/runner… s3744 2 0.004 HSO4- M0 42346.
#> 7 /home/runner… s3744 2 0.004 HSO4- 33S 339.
#> 8 /home/runner… s3744 2 0.004 HSO4- 17O 66.6
#> 9 /home/runner… s3744 2 0.004 HSO4- 34S 2315.
#> 10 /home/runner… s3744 2 0.004 HSO4- 18O 438.
#> # ℹ 6,439 more rows
#> # ℹ 7 more variables: tic <dbl>, it.ms <dbl>, data_group <int>, block <int>,
#> # block_name <chr>, data_type <chr>, segment <int>
# a block that is only added to a specific file
df |> orbi_define_blocks(
start_time.min = 0.2,
end_time.min = 0.8,
in_filename = "ac5"
)
#> ✔ [35ms] orbi_define_blocks() added 1 block to 1 file
#> → block in ac5: covers scans 87 to 344 (0.202 to 0.799 min)
#> # A tibble: 6,449 × 14
#> filepath filename scan.no time.min compound isotopocule ions.incremental
#> <chr> <fct> <int> <dbl> <fct> <fct> <dbl>
#> 1 /home/runner… s3744 1 0.002 HSO4- M0 37803.
#> 2 /home/runner… s3744 1 0.002 HSO4- 33S 326.
#> 3 /home/runner… s3744 1 0.002 HSO4- 17O 61.9
#> 4 /home/runner… s3744 1 0.002 HSO4- 34S 2215.
#> 5 /home/runner… s3744 1 0.002 HSO4- 18O 436.
#> 6 /home/runner… s3744 2 0.004 HSO4- M0 42346.
#> 7 /home/runner… s3744 2 0.004 HSO4- 33S 339.
#> 8 /home/runner… s3744 2 0.004 HSO4- 17O 66.6
#> 9 /home/runner… s3744 2 0.004 HSO4- 34S 2315.
#> 10 /home/runner… s3744 2 0.004 HSO4- 18O 438.
#> # ℹ 6,439 more rows
#> # ℹ 7 more variables: tic <dbl>, it.ms <dbl>, data_group <int>, block <int>,
#> # block_name <chr>, data_type <chr>, segment <int>
