Inspecting Files
You do not always need R or Python to look at a bioinformatics file. A few command line tools check contents, verify integrity and count records. This page collects the useful commands per format, then ends with the checklist to run before any analysis.
General tools
Section titled “General tools”These work on almost any file.
| Command | What it does |
|---|---|
head -n 20 file.txt |
Show the first 20 lines |
tail -n 20 file.txt |
Show the last 20 lines |
wc -l file.txt |
Count lines |
zcat file.txt.gz | head |
Peek at a gzipped file |
file mydata.bam |
Check the actual file type |
ls -lh file.bam |
Check the file size |
The file command earns its place when a file has the wrong extension or arrives
with no documentation.
# View two reads (four lines each) from a compressed FASTQzcat reads.fastq.gz | head -8
# Count reads: divide the line count by 4zcat reads.fastq.gz | wc -lzcat reads.fastq.gz | awk 'END{print NR/4}'
# Check paired-end files agreezcat R1.fastq.gz | wc -lzcat R2.fastq.gz | wc -lIf the two line counts differ, something went wrong in demultiplexing or download.
BAM is binary, so you need samtools.
# Header and first alignmentssamtools view -h file.bam | head -20
# Alignment summary: totals, mapped, paired, duplicatessamtools flagstat file.bam
# Reads per chromosome (needs an index)samtools idxstats file.bam
# Create the index if it is missingsamtools index file.bam
# Depth at specific positionssamtools depth -r chr1:1000000-1001000 file.bamflagstat is the fastest quality check on an alignment.
VCF is text but usually large, compressed with bgzip and indexed with tabix.
# The header: filters, INFO fields, sample namesbcftools view -h file.vcf.gz
# Summary: variant counts, Ts/Tv ratio, quality distributionsbcftools stats file.vcf.gz | head -30
# Extract chromosome, position and alleles per variantbcftools query -f '%CHROM\t%POS\t%REF\t%ALT\n' file.vcf.gz | head
# Count variants, excluding header linesbcftools view -H file.vcf.gz | wc -lBED, GFF and GTF
Section titled “BED, GFF and GTF”These are tab-separated text, so standard tools work.
head file.bed
# Regions per chromosomeawk '{print $1}' file.bed | sort | uniq -c | sort -rn
# Gene count in a GTF: the tabs match the feature column, not a gene namegrep -c " gene " file.gtf
# Every feature type in a GFF/GTFawk -F'\t' '$1 !~ /^#/ {print $3}' file.gtf | sort | uniq -c | sort -rnThe last command reports how many gene, transcript, exon and CDS entries the file carries.
H5AD is binary HDF5. A one-line Python command prints the summary.
python -c "import anndata; adata = anndata.read_h5ad('file.h5ad'); print(adata)"AnnData object with n_obs × n_vars = 5000 × 20000 obs: 'cell_type', 'batch', 'n_counts' var: 'gene_name', 'highly_variable' obsm: 'X_pca', 'X_umap'10x Matrix Market files
Section titled “10x Matrix Market files”The filtered_feature_bc_matrix/ directory holds three files. The dimensions sit
in the third line of the matrix header: rows (genes), columns (cells) and non-zero
entries.
zcat matrix.mtx.gz | head -3zcat barcodes.tsv.gz | wc -lzcat features.tsv.gz | wc -lThe validation checklist
Section titled “The validation checklist”Run these five checks before any analysis.
-
Is the file truncated? Compare the expected size, and test a gzip file with
gzip -t. A BAM of a few kilobytes is not an alignment. -
Is the file sorted? Most BAM tools need coordinate-sorted input. Check with
samtools view -H file.bam | grep SOand look forSO:coordinate. -
Do chromosome names match? Some references write
chr1and others write1, and mixing them fails silently. Comparesamtools idxstats file.bam | headwithbcftools view -h file.vcf.gz | grep contig. -
Is it the right genome build? GRCh37/hg19 and GRCh38/hg38 are not compatible. Chromosome lengths confirm it: chr1 is 249,250,621 bp in hg19 and 248,956,422 bp in hg38.
-
Are paired files consistent? R1 and R2 FASTQ files must hold the same reads in the same order, and BAM/BAI and VCF/TBI index pairs must match.
Summary
Section titled “Summary”head,zcatandwc -lcover text files.samtoolscovers BAM:flagstatfor quality,idxstatsfor per-chromosome counts,view -hfor the header.bcftoolscovers VCF:statsfor summaries,queryfor field extraction.awkandgrepcover BED, GFF and GTF.- A one-line Python command summarises an H5AD file.
- Validate truncation, sorting, chromosome naming, genome build and paired-file consistency before you analyse anything.