Skip to content

Inspecting Files

You do not always need R or Python to look at a bioinformatics file. A few command line tools check contents, verify integrity and count records. This page collects the useful commands per format, then ends with the checklist to run before any analysis.

These work on almost any file.

Command What it does
head -n 20 file.txt Show the first 20 lines
tail -n 20 file.txt Show the last 20 lines
wc -l file.txt Count lines
zcat file.txt.gz | head Peek at a gzipped file
file mydata.bam Check the actual file type
ls -lh file.bam Check the file size

The file command earns its place when a file has the wrong extension or arrives with no documentation.

Terminal window
# View two reads (four lines each) from a compressed FASTQ
zcat reads.fastq.gz | head -8
# Count reads: divide the line count by 4
zcat reads.fastq.gz | wc -l
zcat reads.fastq.gz | awk 'END{print NR/4}'
# Check paired-end files agree
zcat R1.fastq.gz | wc -l
zcat R2.fastq.gz | wc -l

If the two line counts differ, something went wrong in demultiplexing or download.

BAM is binary, so you need samtools.

Terminal window
# Header and first alignments
samtools view -h file.bam | head -20
# Alignment summary: totals, mapped, paired, duplicates
samtools flagstat file.bam
# Reads per chromosome (needs an index)
samtools idxstats file.bam
# Create the index if it is missing
samtools index file.bam
# Depth at specific positions
samtools depth -r chr1:1000000-1001000 file.bam

flagstat is the fastest quality check on an alignment.

VCF is text but usually large, compressed with bgzip and indexed with tabix.

Terminal window
# The header: filters, INFO fields, sample names
bcftools view -h file.vcf.gz
# Summary: variant counts, Ts/Tv ratio, quality distributions
bcftools stats file.vcf.gz | head -30
# Extract chromosome, position and alleles per variant
bcftools query -f '%CHROM\t%POS\t%REF\t%ALT\n' file.vcf.gz | head
# Count variants, excluding header lines
bcftools view -H file.vcf.gz | wc -l

These are tab-separated text, so standard tools work.

Terminal window
head file.bed
# Regions per chromosome
awk '{print $1}' file.bed | sort | uniq -c | sort -rn
# Gene count in a GTF: the tabs match the feature column, not a gene name
grep -c " gene " file.gtf
# Every feature type in a GFF/GTF
awk -F'\t' '$1 !~ /^#/ {print $3}' file.gtf | sort | uniq -c | sort -rn

The last command reports how many gene, transcript, exon and CDS entries the file carries.

H5AD is binary HDF5. A one-line Python command prints the summary.

Terminal window
python -c "import anndata; adata = anndata.read_h5ad('file.h5ad'); print(adata)"
AnnData object with n_obs × n_vars = 5000 × 20000
obs: 'cell_type', 'batch', 'n_counts'
var: 'gene_name', 'highly_variable'
obsm: 'X_pca', 'X_umap'

The filtered_feature_bc_matrix/ directory holds three files. The dimensions sit in the third line of the matrix header: rows (genes), columns (cells) and non-zero entries.

Terminal window
zcat matrix.mtx.gz | head -3
zcat barcodes.tsv.gz | wc -l
zcat features.tsv.gz | wc -l

Run these five checks before any analysis.

  1. Is the file truncated? Compare the expected size, and test a gzip file with gzip -t. A BAM of a few kilobytes is not an alignment.

  2. Is the file sorted? Most BAM tools need coordinate-sorted input. Check with samtools view -H file.bam | grep SO and look for SO:coordinate.

  3. Do chromosome names match? Some references write chr1 and others write 1, and mixing them fails silently. Compare samtools idxstats file.bam | head with bcftools view -h file.vcf.gz | grep contig.

  4. Is it the right genome build? GRCh37/hg19 and GRCh38/hg38 are not compatible. Chromosome lengths confirm it: chr1 is 249,250,621 bp in hg19 and 248,956,422 bp in hg38.

  5. Are paired files consistent? R1 and R2 FASTQ files must hold the same reads in the same order, and BAM/BAI and VCF/TBI index pairs must match.

  • head, zcat and wc -l cover text files.
  • samtools covers BAM: flagstat for quality, idxstats for per-chromosome counts, view -h for the header.
  • bcftools covers VCF: stats for summaries, query for field extraction.
  • awk and grep cover BED, GFF and GTF.
  • A one-line Python command summarises an H5AD file.
  • Validate truncation, sorting, chromosome naming, genome build and paired-file consistency before you analyse anything.